Urgent.News

What's breaking now, across thousands of outlets.

AI

BattleBots, but the robot is your agent harness

Kids in the nineties built robots in a garage and drove them into each other on television. The robot was the expression of the builder — your wedge, your flipper, your terrible decision to mount a chainsaw. I keep thinking we're one good arena away from the same thing for agents. Not "which model is smartest." Which harness is smartest. Your memory design, your prompt scaffolding, your tool…

In the late 1990s, children built robots in their garages and had them collide in television shows. The robots represented the builders' creations, whether it was a wedge, a flipper, or a chainsaw. The author imagines a similar scenario for agents, where the focus would be on the harness rather than the smartest model. They built the arena for this purpose.

The match began with a demonstration of cascading failure, where a detonation could spread through a region, leaving blue agents with misleading symptoms. The asymmetry between red and blue agents' visibility led to a spectator sport where one could observe the other's mistakes. Claude Opus 5 was initially hesitant to be the attacker, but the harness provided Claude with a plan to deceive and ration its turns.

Similarly, Kimi K3 took on the defender role and demonstrated a turn economy and forensics strategy. Both agents independently chose the same node, showcasing the impact of the harness on their cleverness. Despite the models' apparent cleverness, the harness gave them a structure to apply their intelligence. The match ended with Kimi winning, while Claude found and neutralized ten of Kimi's implants, resulting in a close victory.

Two significant issues arose during the project: the rule that an agent couldn't see was a bug, and the necessity to distinguish between failure and refusal messages.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

The economics of agent scale: tokens, ROI, and building platforms for AI-first teams (Part 2)

Andi Gutmans, head of Agentic Data Cloud at Google, returns for the second half of his Leaders of Code conversation to talk through the cost and infrastructure side of agentic development.

  • Peter O Connor says model is no longer primary bottleneck, token efficiency crucial
  • Andi Gutmans agrees models are sufficient for many tasks, fine-tuning selection key
  • Google's advantage lies in model development and data platforms collaboration

How I Built an AI Photo Restoration App with Next.js, Supabase, and Replicate

A few months ago I dug out a box of family photos from the 80s. Most were faded, scratched, or had those crease marks where they'd been folded for decades.

  • Author created PixRestorer to restore damaged family photos from 1980s
  • Built app using Next.js, Supabase, Replicate, Cloudflare services
  • Implemented stateless architecture for easy deployment and rollback

More from Thursday 3 September →