The explanation was right. The policy ID was wrong.
This is a submission for the Kaggle Benchmarking Challenge . What I Benchmarked A support model told me the correct fee, named the applicable policy in its explanation, and then put a different policy in source_ids . A person reading the explanation and software consuming the fields would receive inconsistent guidance. Support Boundary Bench asks whether a model can select the right support…
The explanation in the support model's output provided the correct fee and policy, but the policy ID within the model's source_ids field was incorrect. This inconsistency would lead to different guidance for someone reading the explanation and a software consuming the fields. The benchmark tested two models, OpenAI's GPT-5.4-mini and Google's Gemini-3.7-flash, using the same inputs, prompts, labels, and scoring rules.
The GPT model passed 26 out of 30 pairs, while Gemini passed all 15 pairs. The findings indicate that correct decisions may conceal incorrect evidence, as GPT selected the correct decision type in all cases but passed all structural fields in only 26 out of 30 cases, failing when the policy dates were involved. The replication of the experiment on all cases revealed a similar pattern, where the total score could hide different failures.
Therefore, while a correct decision can be masked by an incorrect policy ID, it is crucial to validate source IDs against product and event date before relying on downstream software.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.