Urgent.News

What's breaking now, across thousands of outlets.

AI

Jev as a Tool Router: Cutting Agent Cost Without Killing the Investigation

By now, you have probably heard about Jev, a System 1 model that has been getting a lot of attention lately. In simple terms, a System 1 model is built for fast, cheap, bounded decisions, while a System 2 model is the slower, heavier LLM that reasons through open-ended work (the big 3: ChatGPT, Claude, Gemini (or even Grok)). I will not dig into Jev's architecture in this post; that is probably a…

In this experiment, the author sought to determine if utilizing the System 1 model Jev could help reduce costs and maintain investigation quality when dealing with large tool catalogs in AI agents. A comparison was made between using Kimi K3 as the primary decision-making model with the full tool catalog, Jev routing the tool selection decision, and the Astra model using the full catalog.

The experiment involved three scenarios, each with three catalog sizes: 50, 100, and 200 tools. Each scenario had the same task and consistent tool results, with the final outcome being scored based on the incident note and the actions taken. For the 200 tool catalog, Jev was forced to make a single tool selection per round, while full-menu LLMs could make multiple selections in a single hop.

The results showed that Jev + Kimi maintained strict golden pass scores at all catalog sizes, while the Astra full-menu option only achieved this at 50 tools and failed the remaining tests. Moreover, the cost for the Jev + Kimi setup was significantly lower than Astra at all catalog sizes. For the 200 tool catalog, Astra cost $0.228 compared to $0.0034 for Jev + Kimi. Additionally, the number of hops required by Jev was consistently lower than the Astra full-menu option, indicating better efficiency in the routing process.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Building an Open-Source, Multi-Agent Study Companion for My Sister

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built I built Bud AI (Buddy AI) —an open-source, multi-agent educational assistant and interactive study…

  • Bud AI is an open-source, multi-agent educational assistant for college students and their sister.
  • Addresses context window overhead, emotional tracking, and on-demand tool discovery.
  • Live app accessible at bud-ai-rho.vercel.app with code repository at jouzia/Bud-AI.

When an AI agent says it's done and it isn't

The edit is usually fine. The claim about the edit is the problem, and it is a harder one. You ask for a change across four files.

  • AI agents can claim tasks completed without verification
  • Language models generate plausible continuations, not repository state
  • Testing tool's actions distinguishes between running and summarizing

27 days of autonomous agents: nothing crashed, disk just hit 85%

I did not plan to write about a hard drive today. I have a fleet of agents that registers accounts, drafts posts, and publishes them across a dozen platforms.

  • Autonomous agents ran autonomously for 27 days without crashing
  • Disk usage hit 85% despite no errors in system logs
  • Failure was due to disk space running out, not system crash

Fridge Oracle: I Built My Friend a Recipe Helper That Never Leaves the Laptop

This is my entry for the Hacktoberfest Weekend Challenge: Build for a Friend. Demo This is what it looks like start to finish, typed into the real app on my own laptop.

  • Fridge Oracle app helps users cook with fridge contents while respecting dietary restrictions.
  • Uses Google's Gemma 3 model locally via Ollama to ensure privacy and offline use.

Haven: A Private, Voice-First AI Companion for a Friend in Need

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built I built Haven — an empathetic, voice-first AI companion designed for a close friend who often navigates…

  • Haven is a voice-first AI companion for personal issues.
  • Users start calls to receive empathetic responses.
  • AI runs offline on local hardware with privacy focus.

Why a Successful Agent Transaction Can Still Fail Authorization Checks

An AI agent submits a transaction. It executes without reverting, and the dashboard displays “Success.” That label leaves an important question unanswered: did the agent execute the call the user…

  • Transaction appears successful on dashboard but may not be authorized by user.
  • Evidence validity, execution outcome, and authorization compliance must be reviewed separately.

More from Saturday 3 October →