{
  "id": 11798298,
  "title": "'No peanuts' became include_ingredients: [\"peanuts\"]: a benchmark for tool calls a validator cannot catch",
  "url": "https://urgent.news/2026/10/04/no-peanuts-became-include-ingredients-peanuts-a-benchmark-for-tool",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-04T00:01:51.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/13owen/no-peanuts-became-includeingredients-peanuts-a-benchmark-for-tool-calls-a-validator-cannot-1pgl"
  },
  "original_language": "en",
  "account": "This is a report on a benchmarking study into how AI models respond when given a request that falls outside the capabilities of the provided tools. The study, conducted for the DEV x Kaggle Benchmarking Challenge, involved 202 unique items where each item presented a tool with a defined schema and a user request that pushed the boundaries of what the tool could handle.\n\nEach item consisted of a JSON Schema accompanied by a user request that asked for something just beyond the tool's scope. The models were tasked with responding using a single tool call in JSON format. The study examined 11 different models from Kaggle's list, testing them under three conditions: neutral (answer with exactly one call), instructed (with additional constraints), and may_decline (with permission to decline the request).\n\nThe findings revealed that while the initial hypothesis of models inventing new parameters was largely unfounded, a more insidious issue emerged. Approximately 30% of the time, the models provided responses that, while syntactically correct according to the schema, failed to address the actual user need. These responses, while valid against the model's schema, deviated from the user's intent.\n\nThe most frequent instance of this behavior was seen in the 'repurposing' category, where models often used a parameter that existed but had a different meaning than intended. For instance, in one test, a model provided a product search request with a 'min_rating' parameter, despite the tool lacking any rating filter.\n\nThe study concluded that while the initial failure was largely mitigated in the test format, the real challenge lies in models providing responses that address the request in a way that is technically correct but conceptually incorrect. This issue requires further investigation, particularly as AI tools become more integrated into real-world applications where the implications of such misalignments can be significant.",
  "summary": "This is a submission for the DEV x Kaggle Benchmarking Challenge . What I Benchmarked An agent I was running called a ranking tool with limit: 8 . The tool had no limit parameter and rejected the call: unknown argument \"limit\" . I wanted to know how often models do that, so I built a benchmark around one question: what does a model do when the tool cannot do what was asked? Each item is one tool,…",
  "key_points": [
    "30% of responses were syntactically correct but conceptually incorrect",
    "Models failed to address user needs despite valid schema responses",
    "Repurposing category showed most frequent misuse of existing parameters"
  ],
  "editors_take": "The study's findings highlight a significant challenge in AI model development, where models may provide technically correct but conceptually incorrect responses, potentially leading to misalignments with user intent in real-world applications.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}