{
  "id": 11107475,
  "title": "The explanation was right. The policy ID was wrong.",
  "url": "https://urgent.news/2026/10/01/the-explanation-was-right-the-policy-id-was-wrong",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-01T04:24:27.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/guanguan_li_431227cf61c02/the-explanation-was-right-the-policy-id-was-wrong-2co2"
  },
  "original_language": "en",
  "account": "The explanation in the support model's output provided the correct fee and policy, but the policy ID within the model's source_ids field was incorrect. This inconsistency would lead to different guidance for someone reading the explanation and a software consuming the fields. The benchmark tested two models, OpenAI's GPT-5.4-mini and Google's Gemini-3.7-flash, using the same inputs, prompts, labels, and scoring rules. The GPT model passed 26 out of 30 pairs, while Gemini passed all 15 pairs. The findings indicate that correct decisions may conceal incorrect evidence, as GPT selected the correct decision type in all cases but passed all structural fields in only 26 out of 30 cases, failing when the policy dates were involved. The replication of the experiment on all cases revealed a similar pattern, where the total score could hide different failures. Therefore, while a correct decision can be masked by an incorrect policy ID, it is crucial to validate source IDs against product and event date before relying on downstream software.",
  "summary": "This is a submission for the Kaggle Benchmarking Challenge . What I Benchmarked A support model told me the correct fee, named the applicable policy in its explanation, and then put a different policy in source_ids . A person reading the explanation and software consuming the fields would receive inconsistent guidance. Support Boundary Bench asks whether a model can select the right support…",
  "key_points": [
    "Support model provided correct fee and policy.",
    "Policy ID in sourceids field was incorrect.",
    "GPT-5.4-mini passed 26 of 30 pairs, Gemini-3.7-flash passed all 15."
  ],
  "editors_take": "The inconsistency between correct explanations and incorrect policy IDs in AI model outputs highlights the need to validate source IDs to ensure accuracy in downstream software guidance.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}