Urgent.News

What's breaking now, across thousands of outlets.

AI

I ran six coding agents on seven local models, 30 times each

Last week I posted a small benchmark on whether coding agents still work when the model you run yourself is shaky at tool calls. It used three runs per task, a handful of models, and a harness I kept private. Fair criticism followed, so here's the bigger, stricter version. Seven models running locally, including the current ones people actually pick today. Six agents. Thirty runs per agent per…

The article details a comprehensive benchmark comparing six coding agents on seven different local models, each tested 30 times. The models used are Qwen3.8-27B, gpt-oss 20B, Devstral Small 2 24B, qwen3-coder 30B, and qwen2.5-coder 7B, 14B, and 32B. Six agents evaluated were Polyglot, pi, goose, Hermes, opencode, and Qwen2.5-coder.

All tests were conducted on a single RTX 5080 16GB GPU with a 32k context. The results show that Polyglot is the only agent that didn't fail on any of the seven models, with its weakest performances on qwen2.5-coder 32B and 7B models. The testing also revealed issues with some models emitting tool calls in plain text instead of through a native channel, affecting the agents' performance.

The article also discusses the impact of prompt size on agent performance and the importance of using a consistent methodology in benchmarking coding agents.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Thursday 1 October →