Urgent.News

What's breaking now, across thousands of outlets.

AI

ButterflyBench: I Changed One Instruction. What Else Did the AI Change?

This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked One of my test prompts said: "Read JSON files instead of CSV." In more than half of my reruns, three models answered by also setting the delimiter and the header-row setting to unspecified . They weren't wrong. A JSON file has no delimiter and no header row. My scoring code was wrong, because it assumed the other 19…

The author of the piece introduces ButterflyBench, a tool designed to measure how much an AI model changes when given a small instruction. The key takeaway is that checking whether a model changes only what you asked it to can be more complex than anticipated, for both the model and the person writing the test. The author ran several AI models through 40 different scenarios, all starting with the same 20-setting specification for a command-line data tool.

The models were tested on various actions like setting values, undoing changes, redoing changes, dropping requirements, or introducing distractors and coupled rules. While some models performed well, others struggled with undoing the earliest change that was still in effect. This highlights the challenges of ensuring an AI model adheres strictly to the given instructions.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Touch Grass Challenge — AI Gives You a Reason to Go Outside

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass What I Built outside | AI plans. You go.

  • Touch Grass Challenge encourages outdoor activities.
  • AI generates personalized monthly outdoor tasks.
  • Local AI inference ensures offline functionality.

AI hackathon: how to test a solution before the final pitch

A team shows a successful model response. A judge changes the input text and gets a different result: a date disappears, an invented city appears or the request hangs.

  • Establish evaluation protocol with testing examples, comparison rules, and post-run data.
  • Provide 30 artificial announcements for development, 12 held out for final evaluation.

More from Wednesday 7 October →