Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems
Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. I argue this measures the wrong axis. The models have saturated accuracy: their mean output lands on the target. What now separates one system from another in practice is precision: how tightly concentrated their outputs are around that target across repeated, identical…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.