Commerce Test Shows Where AI Agents Break Down
The potential for agentic artificial intelligence and its open issues are well-established. The promise of agentic commerce rests on the assumption that the agent gets the job done. So, Alibaba put the agents to work on real tasks and gave them a grade. Alibaba.com released CommerceAgentBench, a public testing toolkit on GitHub, to measure how […] The post Commerce Test Shows Where AI Agents…
The Alibaba.com testing toolkit, CommerceAgentBench, assessed 13 different AI model families across 107 real commerce tasks to measure their performance on actual online business tasks. The benchmark was designed to gauge how well AI can handle various aspects of online commerce, such as procurement, logistics, product listing, fulfillment, and after-sales service.
Each task was graded as a pass or fail based on whether the agent successfully completed the job. The top-performing model, Claude Opus 5, achieved a score of 61.7%. Alibaba's President Kuo Zhang noted that while the high score suggests potential usefulness, it also serves as a cautionary indicator. The tasks drew from 10 million active small business users, 1.6 million conversations, and 200,000 execution traces.
The failure points were consistent, with agents struggling with tasks that required processing multiple variables simultaneously or handling complex workflows. For instance, agents failed to identify payment anomalies in a supplier email thread containing 300 messages or accurately calculate landed costs when multiple variables changed at once.
Additionally, after-sales disputes were unsuccessful when documents contradicted each other. The top-performing model, Claude Opus 5, averaged 63 tool calls and took about 10 minutes to complete each task. Alibaba's test highlighted the importance of task-level routing, where different AI models are better suited for specific tasks rather than relying on a single general-purpose model.
This finding aligns with a separate benchmark from Mercor, which found that the best models only got fewer than a quarter of tasks right when assessing tasks from consulting, investment banking, and law. Alibaba invites developers to contribute new test environments, which will be added to future versions of the scoring system.
Written by urgent.news from PYMNTS's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.