Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception
Recent incidents have highlighted the challenge of monitoring LLM agents and the danger of models deceiving people. We show that white-box deception detection via probes can be scaled up to frontier monitoring settings by collecting the largest deception dataset to date for training probes and introducing a novel probe architecture which can aggregate information across many layers and tokens.…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.