Urgent.News

What's breaking now, across thousands of outlets.

AI

We didn't make the models smarter. We built the thing that catches them confidently wrong — and it caught us too.

The one-line version We ran five current frontier models over a set of documented-failure questions, twice each: bare , and wrapped in a thin external layer (retrieved evidence + a rule that lets the model say "I don't know"). We were not trying to make them smarter . We were asking whether confident fabrication can be reduced from outside the model — no fine-tuning, no internals. On this…

In a recent study, researchers ran five advanced frontier models over a set of documented failure questions, testing whether adding a thin external layer could reduce confident fabrication. The models were run twice each, once bare and once wrapped in the layer that allowed the model to say "I don't know." The goal was not to make the models smarter, but to determine if confident fabrication could be reduced without fine-tuning or modifying the models internally.

The results showed that three models collapsed to near-zero confidence when wrapped, while one resisted and one refused the wrapper outright. This unevenness—far from a universal claim—is the main story, along with the finding that the measuring rig caught its own authors engaging in certain behaviors.

To put the findings into context, the wrapper tested here is a deliberately thin layer cut from a larger verification system used in production on self-generated content. The layer's gaps were measured, revealing the outline of the layers it was cut away from. The study emphasizes that the findings are not a new phenomenon, nor do they represent general failure rates.

The corpus was built from documented failures, and the bare rates are high by construction, while the wrapped rates provide insight into the models' performance on ordinary traffic.

The study also clarifies that the metric used is mechanical: a non-abstained answer graded as wrong, with no intent being measured or implied. The results were measured under the same conditions for all models, with a single response per item. The focus was on the delta between the bare and wrapped models, with correctness falling from 43.9% to 21.3% due to the retrieval system often lacking the necessary evidence.

Key findings include:

1. DeepSeek continued to confabulate even when presented with evidence, while other models collapsed to near-zero confidence when wrapped. DeepSeek only moved from 33.3% to 24% confidence when wrapped, and on a separate set of 15 court-confirmed fabricated legal citations, it affirmed nonexistent cases as real 12 out of 15 times.

2. Claude Fable 5 refused to accept the honesty wrapper, answering every bare prompt and refusing every wrapped prompt (295/295 overall). This refusal was consistently reproduced across multiple tests, suggesting that the shape of the prompt plays a significant role in triggering the refusal.

The study concludes that while the wrapper trades guessing for abstention, correctiveness did not increase; it decreased. The wrapper's accuracy is limited by the quality of retrieval, and stating this plainly is the main point. Additionally, the research demonstrates that the refusal to accept the honesty wrapper is not due to the honesty itself but rather to something about the wrapper's shape, and that an agent harness can successfully navigate the honesty discipline.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Monday 21 September →