Research finds AI agents haven't quite mastered real-world browsing tasks despite claiming they can
Many AI agents lack sufficient safeguards and can't handle multiple tabs very well despite promises.
A new study by Decodo has exposed the shortcomings of AI agents when it comes to real-world browsing tasks, despite vendors' claims of their capabilities. Across 10 different abilities, 45 AI agents were evaluated, but none achieved a perfect score of 20, with Claude for Chrome coming closest at 18 points. The study found that AI agents often fall short in critical areas such as completing transactions and handling third-party integrations.
Claude for Chrome and ChatGPT Chrome Extension were the highest and lowest scoring agents, respectively. Decodo's research highlights that agents typically struggle with transactions, with an average score of just 0.43 out of 2. While these agents can often reach the checkout stage, they struggle to complete purchases. The study also warns that agents may not have adequate safeguards to protect sensitive information like credit card numbers.
Decodo emphasizes that users should match agents to their specific needs rather than relying solely on feature lists.
Written by urgent.news from TechRadar's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.