{
  "id": 12603665,
  "title": "ButterflyBench: I Changed One Instruction. What Else Did the AI Change?",
  "url": "https://urgent.news/2026/10/07/butterflybench-i-changed-one-instruction-what-else-did-the-ai-change",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-07T10:36:14.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/emmimal_alexander_3be8cc7/butterflybench-i-changed-one-instruction-what-else-did-the-ai-change-2chb"
  },
  "original_language": "en",
  "account": "The author of the piece introduces ButterflyBench, a tool designed to measure how much an AI model changes when given a small instruction. The key takeaway is that checking whether a model changes only what you asked it to can be more complex than anticipated, for both the model and the person writing the test. The author ran several AI models through 40 different scenarios, all starting with the same 20-setting specification for a command-line data tool. The models were tested on various actions like setting values, undoing changes, redoing changes, dropping requirements, or introducing distractors and coupled rules. While some models performed well, others struggled with undoing the earliest change that was still in effect. This highlights the challenges of ensuring an AI model adheres strictly to the given instructions.",
  "summary": "This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked One of my test prompts said: \"Read JSON files instead of CSV.\" In more than half of my reruns, three models answered by also setting the delimiter and the header-row setting to unspecified . They weren't wrong. A JSON file has no delimiter and no header row. My scoring code was wrong, because it assumed the other 19…",
  "key_points": [
    "ButterflyBench measures AI model changes with small instructions",
    "Author tests 40 scenarios on 20-setting command-line tool",
    "Some models struggle with undoing earliest changes"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}