Copy article

Google's Research Reveals AI's 70% Factual Accuracy Ceiling

ended 11. December 2025

Google DeepMind just published research that confirms what every burned CTO already suspected: AI gets basic facts wrong 30% of the time, and multimodal AI, the technology reading your invoices, charts, and diagrams—barely manages a coin flip.

The FACTS Benchmark Suite tested 15 leading models across four scenarios measuring whether AI actually tells the truth. Gemini 3 Pro topped the leaderboard at 68.8%. GPT-5 scored 61.8%. Claude 4.5 Opus managed 51.3%. Not one model broke 70% accuracy.

Multimodal AI scored below 47% across the board, the technology enterprises are deploying to extract data from financial reports, interpret medical imaging, and analyse legal documents. The research team needed two automated judges just to verify whether responses were factually accurate, and those judges themselves only achieved 72-78% reliability against human ratings.

Model
 
FACTS Score (Avg)Search (RAG Capability)Multimodal (Vision)
Gemini 3 Pro
 
68.8
 
83.8
 
46.1
 
Gemini 2.5 Pro
 
62.1
 
63.9
 
46.9
 
GPT-5
 
61.8
 
77.7
 
44.1
 
Grok 4
 
53.6
 
75.3
 
25.7
 
Claude 4.5 Opus   
 
51.3
 
73.2
 
39.2
 

Data sourced from the FACTS Team release notes.

Google's own researchers conclude: "All evaluated models achieved an overall accuracy below 70%, leaving considerable headroom for future progress." Translation: if you're building systems that assume AI tells the truth, you're building on quicksand.

The 70% ceiling is empirical proof that human oversight, review loops, and transparent architectures aren't nice-to-haves, they're the only defensible approach to enterprise AI deployment.

We'd like your views:

  • Should factuality benchmarks become mandatory disclosure before enterprise AI procurement?
  • What's the liability exposure when systems scoring 46% accuracy on image interpretation make consequential decisions?
  • Is human-in-the-loop architecture the only defensible approach—or are we just moving liability around?
  • How do you price AI implementation when the technology provider's own research shows 30% error rates?
  • Should there be regulatory accuracy minimums for AI in finance, legal, and medical contexts?
  • What does "production-ready" actually mean when leading models can't crack 70% factual accuracy?
  • Where's the line between AI augmentation that empowers and automation theatre that just shifts blame?

6 responses from the Newspage community

Copy all

Star Quote
Copy

Marketing departments need to be clear about AI accuracy. It's touted as the shiny new business saviour, but it's wrong 30% of the time.

There's a deliberate blending of trigger-based "logical" automation with AI-based judgement that's deeply concerning. Logical automation rarely fails. AI-based judgement fails one time in three.

When reputations, service quality, real outcomes, and at times lives on the line depend on AI decisions, letting it loose without human supervision is simply not acceptable.
Copy

We question other human beings endlessly, yet we blindly trust technology. That’s the paradox we need to confront. Yes, tech companies have brilliant marketers who promise AI will do everything at the click of a button. But marketing has always sold us the idea that without the latest product, our lives fall short. That isn’t new. What is new is our willingness to wholly delegate judgement to machines without hesitation. And then complain when they get things wrong every so often.

There’s nothing wrong with AI being 50-70%. We built it, and most human beings would struggle to be right more than 70% of the time, no matter what they think. The issue isn’t the technology; it’s our belief that it should be flawless. You wouldn’t let even your most capable employee work without any oversight at all. There will be benchmarks and targets in place to monitor them. AI is no different. Human review isn’t a nice-to-have. It’s the scaffolding that keeps the whole system safe.
Copy

If a human was wrong 30% of the time, they’d be on a performance plan, not trusted with mortgages, diagnoses or redundancies – so why are we pretending that’s fine for AI? This isn’t “future of work”, it’s basic responsibility. You cannot dump a 46%-accurate system in the middle of your operation and then act shocked when it ruins real people’s lives. Someone, with an actual job title and a conscience, has to own the final decision. Human-in-the-loop isn’t fluffy HR talk, it’s your only defence. If leaders hide behind “the system says so”, you kill challenge, you silence your brightest people, and you turn accountability into a game of pass the parcel. That’s not innovation. That’s just blame with better branding.
Copy

You cannot build a business on a system that is wrong three times out of ten. Trust takes years to build and only seconds to lose. If you hand your brand over to a machine without checking the work, you are taking a risk that does not make commercial sense. This research proves that expert human oversight is the only way to use these tools safely.
Copy

It is genuinely baffling. In the week Google confirms AI gets basic facts wrong 30% of the time, banking giants announce there will likely be job cuts to rely more on it. This isn't innovation; it’s corporate self-sabotage.

They are voluntarily removing their quality control, their people, just as the data proves the technology isn't ready for unsupervised work. They’re building faster engines while firing the drivers.

This is a golden opportunity for SMEs. While enterprises drown in error management and alienate customers with hallucinating chatbots, agile businesses can take the market. The strategy is simple: use AI for speed, but keep humans for judgment.

Let the big banks race to the bottom on service. In 2026, 'Human Verified' will be the ultimate premium product. The SMEs that combine AI efficiency with human accountability won't just survive this shift; they'll steal the clients the giants are too busy automating to care about.
Copy

Many corporates are applauding AI’s mediocre performance whilst the end user pays the price. They are rolling out an employee into key business areas who, in any real workplace, would be on an improvement plan under close supervision. Humans make mistakes and correct them. AI makes mistakes and cements them, and the public are left to absorb the consequences. A misread payslip becomes a credit block. A flawed pattern match becomes a fraud flag. These systems are nowhere near ready for the high-level decisions they are already making. If AI is going to be responsible for such critical business areas, it must be held to the same standard as any manager: supervise it, challenge it, and never let it make a call that no accountable human would sign off.