Don't Average Pass Rates Across Unequal Token Budgets
You taped two terminal panes to a Friday review. The left pane showed Agent A with eight greens; the right pane showed Agent B with six. Someone asked which one to keep. You almost said A. Then you scrolled the traces. Agent A retried until the cap cried. Agent B stopped after one compile and one test run. Same tasks. Different permission to spend. The ranking was a poster, not a measurement. A…
Don't average pass rates across unequal token budgets. A pass rate only indicates if tests completed successfully, without revealing token usage or budget constraints. Comparing agents with different budget allowances fails to provide a true benchmark. The article proposes a protocol to record detailed metrics like pass rate, truncation rate, and tokens-per-success, rather than relying on a single percentage metric.
It emphasizes the importance of declaring budgets upfront, logging every attempt, and publishing three key numbers together to accurately compare agents.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.