Urgent.News

What's breaking now, across thousands of outlets.

AI

Passing Once Isn't Reliable — This Week's Agent Engineering Puts the Harness Before the Model

This digest covers AI agent developments from 2026-08-18 to 2026-08-25: orchestration patterns, tool/function calling, memory, planning loops, multi-agent coordination, and agent evaluation. 🔥 Highlights AgentWeave: Routing Before Reasoning for Efficient Function Calling in Tool-Rich Language Models — cuts tool exposure 70%, latency 51%. One Success Isn't Reliability: Thinkingbox, a Sandbox and…

This week's AI agent developments focus on orchestration patterns, tool/function calling, memory, planning loops, multi-agent coordination, and agent evaluation. AgentWeave, a new system, filters tool sets before prompts are processed, reducing tool exposure by 70%, input tokens by 62%, and latency by 51%. However, pass@1 success rates in Thinkingbox, a sandbox and benchmark, do not guarantee reliability as seen in the drop from 65.36% pass@1 to 25.25% pass@20.

Spine-Branch Coordination for Multi-agent Computer Use tackles VM state merging in multi-agent systems, lifting success rate by 6-16.5 points and cutting cost per task by 34-70%. AutoSaddler optimizes the agent's external harness using execution traces, resulting in gains of 9-10 points across three benchmarks. LangSmith Tuned Evaluators, starting with Perceived Error, grade 100% of production traffic, cutting evaluation cost by 82% while maintaining quality.

Anthropic Engineering Blog and LangChain / LangGraph Blog introduced new tools and frameworks for testing and evaluating agents, emphasizing the importance of the agent harness and its impact on performance.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Why LLM reasoning isn't enough for medical scheduling math

I’ve seen plenty of people try to make Claude or GPT-4 act like a specialized scheduler. They prompt it heavily: "You are a precise medical assistant.

  • LLMs excel at simple scheduling but struggle with complex constraints
  • Probabilistic nature of LLMs leads to hallucinations in deterministic scheduling
  • Injection Day Alignment MCP server offers structured tools for medical scheduling

Jak rozjet lokální LLM napojený na Home Assistant

Když se nedaří rozjet lokální model, bývá to většinou o jedné špatné komponentě. Uživatel z r/LocalLLaMA se o tom přesvědčil na vlastní kůži: po roce vzdávání stačila hodina práce s llama.cpp a Qwen…

  • User overcame challenge of launching local LLM for Home Assistant integration
  • Qwen 3.8 27B model connected to Home Assistant server and customized with prompts and screenshots
  • Community provided valuable tips to help user set up local LLM successfully

Why Your OpenAPI Spec Isn't Enough for AI Agents

OpenAPI describes an API. Agent Readiness describes whether an agent can actually use it. Your API has a complete OpenAPI spec. Every endpoint, schema, and response code is documented.

  • OpenAPI spec describes API endpoints and schemas
  • OpenAPI alone insufficient for AI agents
  • Agent Readiness provides additional context

More from Tuesday 25 August →