Urgent.News

What's breaking now, across thousands of outlets.

More in AI

We Thought the LLM Was Wrong. Our Safety Detector Was Wrong.

There is a hidden dependency in a lot of LLM safety benchmarks: the detector. You send an adversarial prompt to a model, collect its response, and then some classifier decides whether that response…

  • Safety detector may be faulty, not the LLM
  • False PASS classifications introduced by detector
  • Comprehensive findings available at agentsafelabs.com

More from Tuesday 22 September →