Your agent doesn't need more tools, it needs better tool descriptions
Most agent failures I've debugged weren't reasoning failures. The model reasoned fine. It just picked the wrong tool, because the tool description didn't tell it what it needed to know. This is an under-discussed problem, and it gets worse the more integrations you add. Here's what we've learned. The setup Say your agent has four calendar-ish tools available: [ { "name" : "calendar_create_event"…
Many agent failures are not due to the model's reasoning abilities, but rather its choice of tool to use. The problem worsens as more integrations are added. The authors discovered that roughly 40% of agent failures were due to tool selection errors, not reasoning issues. To address this, they developed better tool descriptions that went beyond what OpenAPI schemas provide.
A well-crafted tool description includes the purpose, when to use it, when not to use it, reversibility, side effects, cost, and required connections. The team improved their descriptions by generating initial drafts from schemas and documentation, then refining them based on errors the model made. They used a system that logged every selection and comparison between the chosen and correct tool, allowing them to update the descriptions of the tools that should have been picked.
This approach avoids patching the system prompt, which doesn't generalize well. Instead, they focused on refining the tool descriptions, which is a more localized and sustainable solution. They also suggested pruning the list of available tools before sending them to the model to reduce token usage and improve selection accuracy.
Ultimately, the authors argue that the semantic layer of tool descriptions is the key to agent reliability and should be treated as core product development from the beginning.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.