Urgent.News

What's breaking now, across thousands of outlets.

AI

AI could be costing you money: new study finds chatbots get most financial questions wrong

According to a new study conducted by an AI and technology firm, AI tools are more susceptible to giving out false information in regards to its users’ financial inquiries.

AI could be costing you money: new study finds chatbots get most financial questions wrong

A recent study has revealed that AI-powered chatbots are frequently providing incorrect financial advice, which could lead to substantial financial losses for users. According to the findings, 66% of Americans who have used generative AI (GenAI) have turned to it for financial advice, with the figure rising to 82% among Gen Z and Millennials. Finance is the second most common use case for GenAI, trailing only health and wellness.

However, the reliability of AI models in providing accurate financial guidance is questionable. A study by Saturn, an AI and technology firm, tested 18 popular AI tools, including ChatGPT, Gemini, Claude, and Copilot. The results indicated that, on average, these models were only 43% accurate when answering financial questions, meaning they were wrong 57% of the time.

When presented with complex, multi-step scenarios that required precise financial figures and tax rules, the accuracy rate plummeted to just 12%, with incorrect answers in 88% of the cases.

The study found that free AI models performed significantly worse than paid models, with free models having a 63% failure rate compared to 49% for paid models. Claude Haiku 4.5 was the worst-performing free model, producing incorrect or incomplete answers 82% of the time, while ChatGPT-5.6 Luna (max) was the best-performing free model, although it still had a 56% error rate. Anthropic's paid Claude Opus 5 model, while the best among the tested, still failed to provide correct answers in almost 40% of cases.

Furthermore, the investigation uncovered specific errors that could result in severe financial consequences. For example, one model suggested that a college graduate could stop paying student loans by moving abroad, when in fact, this could lead to higher monthly repayments. Another tested model incorrectly informed a borrower that taking a mortgage payment holiday would not impact their credit score.

The implications of these findings are significant, as many individuals are increasingly relying on AI assistants for financial decisions. A March 2026 report from EY showed that 53% of respondents prefer using AI to make financial investment decisions, while 14% favor autonomous AI, and 33% would not use any AI in such cases. Saturn's study suggests that this preference for AI financial advice may decrease as users become more aware of the model's unreliability and potential for causing significant financial harm.

Written by urgent.news from Tom's Guide's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at tomsguide.com →

More in AI

Reducing medical claims review time with AI on AWS: The EXL Medical IDP solution

EXL built an AI-powered Medical intelligent document processing (IDP) solution on AWS, combining IDP with domain-specific large language models on Amazon SageMaker and Amazon Bedrock to extract…

  • EXL Medical IDP solution reduces claims review time with AI on AWS
  • AI-powered Medical intelligent document processing merges IDP with domain-specific LLMs
  • Solution automates document ingestion, classification, extraction, enrichment, and output delivery

How Benchling secured multi-tenant AI agents with Amazon Bedrock AgentCore

Learn how Benchling built a defense-in-depth security architecture to run untrusted, AI agent-generated scientific code across thousands of life sciences tenants using Amazon Bedrock AgentCore Code…

  • Benchling deployed multi-tenant AI agents with Amazon Bedrock AgentCore
  • Implemented defense-in-depth security architecture using AWS account isolation
  • VPC mode with AgentCore Code Interpreter blocked unauthorized network vectors

More from Monday 21 September →