Google's Research Reveals AI's 70% Factual Accuracy Ceiling
Google DeepMind just published research that confirms what every burned CTO already suspected: AI gets basic facts wrong 30% of the time, and multimodal AI, the technology reading your invoices, charts, and diagrams—barely manages a coin flip.
The FACTS Benchmark Suite tested 15 leading models across four scenarios measuring whether AI actually tells the truth. Gemini 3 Pro topped the leaderboard at 68.8%. GPT-5 scored 61.8%. Claude 4.5 Opus managed 51.3%. Not one model broke 70% accuracy.
Multimodal AI scored below 47% across the board, the technology enterprises are deploying to extract data from financial reports, interpret medical imaging, and analyse legal documents. The research team needed two automated judges just to verify whether responses were factually accurate, and those judges themselves only achieved 72-78% reliability against human ratings.
| Model | FACTS Score (Avg) | Search (RAG Capability) | Multimodal (Vision) |
| Gemini 3 Pro | 68.8 | 83.8 | 46.1 |
| Gemini 2.5 Pro | 62.1 | 63.9 | 46.9 |
| GPT-5 | 61.8 | 77.7 | 44.1 |
| Grok 4 | 53.6 | 75.3 | 25.7 |
| Claude 4.5 Opus | 51.3 | 73.2 | 39.2 |
Data sourced from the FACTS Team release notes.
Google's own researchers conclude: "All evaluated models achieved an overall accuracy below 70%, leaving considerable headroom for future progress." Translation: if you're building systems that assume AI tells the truth, you're building on quicksand.
The 70% ceiling is empirical proof that human oversight, review loops, and transparent architectures aren't nice-to-haves, they're the only defensible approach to enterprise AI deployment.
We'd like your views:
- Should factuality benchmarks become mandatory disclosure before enterprise AI procurement?
- What's the liability exposure when systems scoring 46% accuracy on image interpretation make consequential decisions?
- Is human-in-the-loop architecture the only defensible approach—or are we just moving liability around?
- How do you price AI implementation when the technology provider's own research shows 30% error rates?
- Should there be regulatory accuracy minimums for AI in finance, legal, and medical contexts?
- What does "production-ready" actually mean when leading models can't crack 70% factual accuracy?
- Where's the line between AI augmentation that empowers and automation theatre that just shifts blame?





