Urgent.News

What's breaking now, across thousands of outlets.

World

My factual-recall tasks were scoring format, not facts

Originally published at erikhill.dev . The numbers below are checked against the repository they come from. My factual-recall tasks were scoring format, not facts This is a finding about my own harness. The suspect is the probe, not the models it measures. What happened I built a detector that decides whether a day's drift run contains anything worth writing up. The first thing it did was accuse…

During the investigation of a drift detection tool, it was discovered that the "fact-element" task, which tests a model's ability to recall specific information like the chemical symbol for gold, was scoring based on formatting rather than actual facts. Claude Sonnet 5 failed this task 60 times out of 60, while GPT-4o mini and Llama 3.1 8B failed it 59 and 23 times respectively.

The issue lies in how the models interpret instructions - some interpret them as constraints they must obey, while others see them as hints about what is being asked. This led to GPT-4o mini incorrectly answering "Mercury" to the question about gold's chemical symbol, being recorded wrong 59 times out of 59. While the exact cause and solution are still unclear, it is suggested that these tasks be retagged as instruction-following instead of factual-recall to avoid conflating formatting issues with knowledge recall.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in World

More from Wednesday 23 September →