Saturn's 10,000-Answer Test Finds AI Fails 57% of Money Queries — Bittensor (TAO) in Focus

Saturn's test of 18 AI models found a 57% failure rate on personal finance questions, rising to 88% on hard queries — while trust in AI advice keeps climbing.

(03:53 PM UTC)
4 min read
AI SummaryAI
  • Saturn's test found mainstream AI models failed 57% of personal finance questions.
  • Failure rates reached 88% on harder multi-step queries across 18 tested models.
  • Free AI models failed 63% of answers versus 49% for paid versions.
  • Claude Opus 5 in reasoning mode led the field yet failed 39% of answers.
k7rq2fdm

10,000 Answers, 57% Wrong

The verdict is in, and it is uncomfortable reading for anyone who has asked a chatbot where to put their savings: mainstream AI models delivered wrong or incomplete answers to 57% of personal-finance questions in a test compiled by the British fintech firm Saturn, and the failure rate climbed to 88% on harder, multi-step queries — the kind real households actually face. The exercise ran 121 questions through 18 free and paid models from providers including ChatGPT, Alphabet's Gemini, Claude and Copilot, repeating each question up to five times to probe consistency and collecting more than 10,000 answers in total. An answer was scored as a failure when it contained a factual error, skipped something material, or left out a required warning. Free tiers fared worst, failing 63% of responses against 49% for paid versions; on the hardest questions, free models were wrong 93% of the time. Claude Opus 5 in reasoning mode topped the table, yet still missed on 39% of answers, with errors spanning bad arithmetic, overlooked tax changes and rules that simply do not exist. One answer on pension tax could have exposed a saver to a £17,500 charge from HM Revenue and Customs. “Millions of people are trusting the AI models for money advice, but they are getting wrong answers that can lose them money,” said Amal Jolly, Saturn's chief executive. The timing sharpens the point: consumer reliance on chatbots for money decisions has grown sharply, which turns a 57% error rate from a lab curiosity into a live consumer-protection problem — and one that reaches directly into crypto, where self-directed investors routinely lean on AI tools with no regulated advice layer behind them.

Trust Keeps Climbing Anyway

The sharpest tension in the data is that confidence in these tools keeps rising even as their accuracy flounders. An EY survey spanning 18,000 consumers worldwide found that 49% had already used AI to support savings and investment decisions. Britain's financial regulator reported in August that four in five less-experienced investors had turned to AI for help with investing, and 56% of those surveyed said they trust the tools — a higher share than said the same of television and radio, at 47%. The same research flagged a more worrying misconception: 44% of respondents wrongly believe AI-generated financial information is regulated. A separate PensionBee survey of 1,000 US adults found nearly six in ten would act on money guidance without independently checking it, and almost one in four said a chatbot had already given them wrong information about their finances. Exposure also no longer requires a desktop browser: assistants are embedded at the operating-system level, from Apple's Siri to Gemini on Android, so every age group is in scope. Jolly's structural point stands: AI financial advice sits outside the regulatory perimeter, leaving users without the compensation rights that come with a human adviser, and he has urged the FCA to act quickly. Readers tracking the market in real time can follow live spot and futures prices on Binance.

Bittensor (TAO) and the Trust Gap

COINOTAG's read is that the Saturn results land hardest on the corner of crypto that markets itself on machine intelligence. Bittensor (TAO), the decentralized network that pays contributors for machine-learning models, trades on the same credibility curve the test just dented — a retail public that discovers chatbots fail 57% of money questions may discount AI-token narratives alongside closed-source assistants. The deeper issue is verification: a blockchain oracle exists precisely because unreliable inputs break automated money, and financial chatbots carry no equivalent check. Until model reliability improves and the advice gap is regulated, decentralized AI remains a trust story first and a compute story second.

COINOTAG News Desk

COINOTAG News Desk

COINOTAG's editorial and research desk.

How our News Desk works
AI-Assisted

AI-generated, AI-reviewed, under COINOTAG editorial oversight.