Jev as a Tool Router: Cutting Agent Cost Without Killing the Investigation
By now, you have probably heard about Jev, a System 1 model that has been getting a lot of attention lately. In simple terms, a System 1 model is built for fast, cheap, bounded decisions, while a System 2 model is the slower, heavier LLM that reasons through open-ended work (the big 3: ChatGPT, Claude, Gemini (or even Grok)). I will not dig into Jev's architecture in this post; that is probably a…
In this experiment, the author sought to determine if utilizing the System 1 model Jev could help reduce costs and maintain investigation quality when dealing with large tool catalogs in AI agents. A comparison was made between using Kimi K3 as the primary decision-making model with the full tool catalog, Jev routing the tool selection decision, and the Astra model using the full catalog.
The experiment involved three scenarios, each with three catalog sizes: 50, 100, and 200 tools. Each scenario had the same task and consistent tool results, with the final outcome being scored based on the incident note and the actions taken. For the 200 tool catalog, Jev was forced to make a single tool selection per round, while full-menu LLMs could make multiple selections in a single hop.
The results showed that Jev + Kimi maintained strict golden pass scores at all catalog sizes, while the Astra full-menu option only achieved this at 50 tools and failed the remaining tests. Moreover, the cost for the Jev + Kimi setup was significantly lower than Astra at all catalog sizes. For the 200 tool catalog, Astra cost $0.228 compared to $0.0034 for Jev + Kimi. Additionally, the number of hops required by Jev was consistently lower than the Astra full-menu option, indicating better efficiency in the routing process.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.