WebArena, GAIA and Agentic Benchmarks
An agentic benchmark does not compare a string to an answer key. It puts a system into an environment, lets it act, and then inspects the environment. That is a much better measurement of whether something worked — and it makes the resulting number depend on a great deal more than the model. What makes an agentic benchmark different On MMLU or HumanEval the system produces one artefact and…
We haven't written up this one. Dev.to has the full story — the link below goes straight to it.


