{
  "id": 10990867,
  "title": "What agent frameworks cost on the wire: measurements from agentic-arena",
  "url": "https://urgent.news/2026/09/30/what-agent-frameworks-cost-on-the-wire-measurements-from-agentic-arena",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-30T17:05:46.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/code-with-rashid/what-agent-frameworks-cost-on-the-wire-measurements-from-agentic-arena-13o7"
  },
  "original_language": "en",
  "account": "Framework comparisons often focus on feature lists, but agentic-arena provides a more objective measurement approach. It fixes the model, tools, datasets and iteration budget, then quantifies the differences between various frameworks. These measurements are taken against a scripted (mock) model, ensuring consistent turn-by-turn results across all frameworks. This allows us to focus solely on wire and behavior measurements, without assessing the quality of the generated answers.\n\nOne key metric is the number of prompt tokens per item in a 15-item tool-use task. The vanilla (stdlib loop) baseline uses 753.5 prompt tokens. All other frameworks match this baseline at a 1.00x multiplier, except for smolagents, which uses 3,295.5 tokens - a 3.90x increase. The smolagents framework sends a 3,824-token system prompt that includes twice the tools it has already provided as a schema, resulting in a substantial inefficiency.\n\nAnother crucial factor is how frameworks handle provider failures. The mock provider was scripted to return 429, 500 and 400 status codes. The hand-rolled baseline framework has no retry mechanism, so a single 429 error results in losing an item. However, all frameworks can survive at least one 429 error. The smolagents framework is the only one that can handle three consecutive 429 errors, by sleeping for 2-4 minutes on each affected item. While this keeps the individual item score unaffected, the overall batch throughput is impacted.",
  "summary": "Most framework comparisons argue from feature lists. I wanted numbers, so agentic-arena holds the model, tools, datasets and iteration budget fixed and measures what each framework does differently. Read this first: everything below is measured against a scripted (mock) model, so the turns are byte-identical for every framework. That makes these wire and behaviour measurements. They say nothing…",
  "key_points": [
    "Vanilla baseline uses 753.5 prompt tokens per item in a 15-item tool-use task",
    "smolagents framework uses 3,295.5 tokens, a 3.90x increase over vanilla baseline",
    "smolagents is the only framework that can handle three consecutive 429 errors"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}