Urgent.News

What's breaking now, across thousands of outlets.

AI

Alibaba President: AI agents can talk, but can they actually do the work?

The strongest frontier model we tested successfully completed 61.7% of the tasks — high enough to be useful and low enough to be a warning.

Alibaba President: AI agents can talk, but can they actually do the work?

Alibaba's President expresses doubt about whether AI agents can truly perform commercial work, despite their ability to engage in conversation. He explains that AI requires a different approach than simply measuring intelligence through benchmarks, as business work involves complex tasks like sourcing, marketing, web development, customer support, and handling returns.

The Accio team at Alibaba.com developed a test called CommerceAgentBench, which evaluates AI agents' performance across 107 end-to-end tasks in real e-commerce operations. The test covers procurement, logistics, product listing, fulfillment, and after-sales service. While the strongest AI model successfully completed 61.7% of the tasks, it still fell short in several areas, such as spotting payment anomalies and handling multi-leg shipping routes.

The ranking of AI models varied by task category, indicating that the choice of model should depend on the specific commercial task at hand. The benchmark test helps businesses determine which workflows can be automated and which still require human oversight, enabling precision delegation.

Written by urgent.news from Fortune's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at fortune.com →

More in AI

More from Wednesday 9 September →