Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark
Safety benchmarks typically report one overall score for a suite of datasets, each of which may target one or more safety-related attributes, so models with similar overall scores can have very different attribute profiles. Comparing models is more tractable at the level of individual attributes, yet it is often unclear whether even a single dataset's scores isolate any single attribute. One…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.