{
  "id": 11226929,
  "title": "I ran six coding agents on seven local models, 30 times each",
  "url": "https://urgent.news/2026/10/01/i-ran-six-coding-agents-on-seven-local-models-30-times-each",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-01T16:00:55.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/gsirigu/i-ran-six-coding-agents-on-seven-local-models-30-times-each-2m82"
  },
  "original_language": "en",
  "account": "The article details a comprehensive benchmark comparing six coding agents on seven different local models, each tested 30 times. The models used are Qwen3.8-27B, gpt-oss 20B, Devstral Small 2 24B, qwen3-coder 30B, and qwen2.5-coder 7B, 14B, and 32B. Six agents evaluated were Polyglot, pi, goose, Hermes, opencode, and Qwen2.5-coder. All tests were conducted on a single RTX 5080 16GB GPU with a 32k context. The results show that Polyglot is the only agent that didn't fail on any of the seven models, with its weakest performances on qwen2.5-coder 32B and 7B models. The testing also revealed issues with some models emitting tool calls in plain text instead of through a native channel, affecting the agents' performance. The article also discusses the impact of prompt size on agent performance and the importance of using a consistent methodology in benchmarking coding agents.",
  "summary": "Last week I posted a small benchmark on whether coding agents still work when the model you run yourself is shaky at tool calls. It used three runs per task, a handful of models, and a harness I kept private. Fair criticism followed, so here's the bigger, stricter version. Seven models running locally, including the current ones people actually pick today. Six agents. Thirty runs per agent per…",
  "key_points": [
    "Six coding agents tested on seven local models",
    "Tests conducted 30 times each on RTX 5080 GPU",
    "Polyglot agent showed consistent performance"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}