Urgent.News

What's breaking now, across thousands of outlets.

AI

What agent frameworks cost on the wire: measurements from agentic-arena

Most framework comparisons argue from feature lists. I wanted numbers, so agentic-arena holds the model, tools, datasets and iteration budget fixed and measures what each framework does differently. Read this first: everything below is measured against a scripted (mock) model, so the turns are byte-identical for every framework. That makes these wire and behaviour measurements. They say nothing…

Framework comparisons often focus on feature lists, but agentic-arena provides a more objective measurement approach. It fixes the model, tools, datasets and iteration budget, then quantifies the differences between various frameworks. These measurements are taken against a scripted (mock) model, ensuring consistent turn-by-turn results across all frameworks. This allows us to focus solely on wire and behavior measurements, without assessing the quality of the generated answers.

One key metric is the number of prompt tokens per item in a 15-item tool-use task. The vanilla (stdlib loop) baseline uses 753.5 prompt tokens. All other frameworks match this baseline at a 1.00x multiplier, except for smolagents, which uses 3,295.5 tokens - a 3.90x increase. The smolagents framework sends a 3,824-token system prompt that includes twice the tools it has already provided as a schema, resulting in a substantial inefficiency.

Another crucial factor is how frameworks handle provider failures. The mock provider was scripted to return 429, 500 and 400 status codes. The hand-rolled baseline framework has no retry mechanism, so a single 429 error results in losing an item. However, all frameworks can survive at least one 429 error. The smolagents framework is the only one that can handle three consecutive 429 errors, by sleeping for 2-4 minutes on each affected item. While this keeps the individual item score unaffected, the overall batch throughput is impacted.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Wednesday 30 September →