Cents Matter: Does a Model Notice When One Financial Fact Changes?
This is a submission for the Kaggle Benchmarking Challenge . What I Benchmarked A financial check can fail while looking almost right. A model might recognize an error but give the wrong correction, round at the wrong stage, or confidently decide that two identical amounts are duplicate payments without the evidence to support that conclusion. I have a degree in Accounting, and I wanted a small,…
Three distinct AI models were evaluated using the "Cents Matter" benchmark, which consists of 40 synthetic financial cases. The cases test a model's ability to correctly handle monetary reasoning and accounting principles, even when a single fact changes. Each case provides three key fields: status, expected cents, and discrepancy cents. The goal is for the model to correctly identify and report all three fields when given a case.
The findings show that Gemini 3.7 Flash achieved the highest accuracy with 39 correct cases out of 40 (97.5%). gpt-oss-20b came in second with 38/40 correct (95%), while Claude Haiku performed the poorest, correctly answering only 19/40 cases (47.5%). It's crucial to note that the models were assessed based solely on the parsed financial decision, not raw JSON data.
Some key observations include:
1. Models sometimes provide the correct warning but the wrong correction. For example, Gemini 3.7 Flash recognized an error in a cashflow scenario but incorrectly calculated the discrepancy.
2. Paired scores reveal cases where a model passes individual cases but fails as a pair. Claude Haiku passed 19 individual cases but only two complete pairs, while Gemini 3.7 Flash and gpt-oss-20b performed better in this regard.
3. Consistency across all three output fields is not a guarantee of correctness. Claude Haiku produced inconsistent answers in three instances, even when the status was marked as "ok".
The benchmark emphasizes the importance of cents-level financial reasoning and highlights that a model's performance can vary significantly based on its ability to handle individual facts and their consequences within a financial scenario.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.