{
  "id": 9441165,
  "title": "My factual-recall tasks were scoring format, not facts",
  "url": "https://urgent.news/2026/09/23/my-factual-recall-tasks-were-scoring-format-not-facts",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-23T23:15:35.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/agentdev9/my-factual-recall-tasks-were-scoring-format-not-facts-j4m"
  },
  "original_language": "en",
  "account": "During the investigation of a drift detection tool, it was discovered that the \"fact-element\" task, which tests a model's ability to recall specific information like the chemical symbol for gold, was scoring based on formatting rather than actual facts. Claude Sonnet 5 failed this task 60 times out of 60, while GPT-4o mini and Llama 3.1 8B failed it 59 and 23 times respectively. The issue lies in how the models interpret instructions - some interpret them as constraints they must obey, while others see them as hints about what is being asked. This led to GPT-4o mini incorrectly answering \"Mercury\" to the question about gold's chemical symbol, being recorded wrong 59 times out of 59. While the exact cause and solution are still unclear, it is suggested that these tasks be retagged as instruction-following instead of factual-recall to avoid conflating formatting issues with knowledge recall.",
  "summary": "Originally published at erikhill.dev . The numbers below are checked against the repository they come from. My factual-recall tasks were scoring format, not facts This is a finding about my own harness. The suspect is the probe, not the models it measures. What happened I built a detector that decides whether a day's drift run contains anything worth writing up. The first thing it did was accuse…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}