Urgent.News

What's breaking now, across thousands of outlets.

AI

'No peanuts' became include_ingredients: ["peanuts"]: a benchmark for tool calls a validator cannot catch

This is a submission for the DEV x Kaggle Benchmarking Challenge . What I Benchmarked An agent I was running called a ranking tool with limit: 8 . The tool had no limit parameter and rejected the call: unknown argument "limit" . I wanted to know how often models do that, so I built a benchmark around one question: what does a model do when the tool cannot do what was asked? Each item is one tool,…

This is a report on a benchmarking study into how AI models respond when given a request that falls outside the capabilities of the provided tools. The study, conducted for the DEV x Kaggle Benchmarking Challenge, involved 202 unique items where each item presented a tool with a defined schema and a user request that pushed the boundaries of what the tool could handle.

Each item consisted of a JSON Schema accompanied by a user request that asked for something just beyond the tool's scope. The models were tasked with responding using a single tool call in JSON format. The study examined 11 different models from Kaggle's list, testing them under three conditions: neutral (answer with exactly one call), instructed (with additional constraints), and may_decline (with permission to decline the request).

The findings revealed that while the initial hypothesis of models inventing new parameters was largely unfounded, a more insidious issue emerged. Approximately 30% of the time, the models provided responses that, while syntactically correct according to the schema, failed to address the actual user need. These responses, while valid against the model's schema, deviated from the user's intent.

The most frequent instance of this behavior was seen in the 'repurposing' category, where models often used a parameter that existed but had a different meaning than intended. For instance, in one test, a model provided a product search request with a 'min_rating' parameter, despite the tool lacking any rating filter.

The study concluded that while the initial failure was largely mitigated in the test format, the real challenge lies in models providing responses that address the request in a way that is technically correct but conceptually incorrect. This issue requires further investigation, particularly as AI tools become more integrated into real-world applications where the implications of such misalignments can be significant.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Why 'Local-First' Is the New Stack: How to Build Data-Sovereign AI Apps with Local LLMs, Private Vectors, and Zero-Cloud Dependencies

Originally published on tamiz.pro . For the past decade, the default architecture for software development has been heavily skewed toward the cloud.

  • Local-First architecture ensures data stays within user control.
  • Advances in local AI technologies enable LLMs to run on consumer hardware.
  • Practical implementation uses Python, Ollama, and ChromaDB for zero-cloud AI systems.

More from Sunday 4 October →