Urgent.News

What's breaking now, across thousands of outlets.

AI

The underwhelming results of AI performance metrics should surprise exactly no one

Somebody at Meta built an internal leaderboard called Claudeonomics . It ranked the company’s top 250 consumers of AI tokens and handed out titles like “Token Legend” and “Cache Wizard.” Engineers set agents running for hours to climb it. In one thirty-day stretch, employees on the dashboard burned through more than 60 trillion tokens. Meanwhile, a Disney employee reportedly interacted with an AI…

The underwhelming results of AI performance metrics should surprise exactly no one

Meta’s internal leaderboard, Claudeonomics, was created to rank the company’s top 250 AI consumers and awarded titles such as “Token Legend” and “Cache Wizard.” Engineers spent hours running agents to climb the leaderboard, consuming over 60 trillion AI tokens in a 30-day period. Meanwhile, a Disney employee reportedly interacted with an AI assistant 460,000 times in just nine days.

JPMorgan, KPMG, Amazon, and Accenture have reportedly been tracking employee AI activity, even incorporating "AI-driven impact" into formal performance reviews. This behavior is known as "tokenmaxxing," where people learned how to artificially inflate their AI metrics without necessarily gaining real value.

The results of these AI performance metrics should come as no surprise. Past experiences have shown that any measure with consequences is prone to gaming, as stated by Goodhart's Law. When a metric becomes a target, people respond by manipulating the numbers to meet the target. In this case, employees used work that didn't require AI or employed agents and digital delegates to execute processes just to boost their token usage.

The leaderboard was shut down once it went public, and Amazon reportedly shut down its own AI performance dashboard, with some observers suggesting it incentivized employees to cheat.

There are five key lessons from the history of performance measurements that apply to AI metrics as well. Firstly, any measure with consequences will be gamed, because people respond to instructions. Secondly, when a metric is gamed, its predictive power decreases, making it less useful for forecasting outcomes. Thirdly, inputs are not the same as results.

Tokens are an input, and measuring outcomes separately is crucial. Many businesses learned this lesson the hard way, such as in call centers and hospitals. Fourthly, what's measurable can crowd out what truly matters, as illustrated by Robert McNamara's attempt to manage the Vietnam War using statistical indicators. Lastly, token usage, while easy to measure, does not necessarily indicate the real value created by AI. It's essential to focus on the outcomes that matter most, even if they are difficult to quantify.

Written by urgent.news from Fast Company's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at fastcompany.com →

More in AI

More from Wednesday 9 September →