How We Caught an LLM Scoring Bug and Fixed It Without Compromising Our Research
A 42-of-42 failure rate in our LLM benchmark was a parser bug. How we found it, fixed our scorer without fudging results, and a checklist for your evals.
An analysis of a large language model (LLM) benchmark revealed a scoring bug that led to a 100% failure rate in one cell. Upon closer inspection, the issue was traced back to the parser, which had two main problems. The first issue was that the model did not repeat a required field, as our code already knew the project ID. The second issue was that the parser had strict requirements, including the model repeating a field we could fill in ourselves, and an exact JSON object anchored to the end of the response.
These format issues were present across multiple models and prompting strategies, but the Llama-3.3-70B model experienced the worst of them when using the structured prompt. To fix the scorer without compromising the research, the team wrote down the rules before making any changes. These rules included never regenerating, fixing format, not touching the content, leaving the schema alone, checking what was left, and disclosing the bug and fix in the paper.
After implementing these changes, the total parse failures across the grid decreased from 88 to 49, with the Llama-3.3-70B model's recall improving from 0.0 to around 0.43 on the structured prompt.
Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.