Copy article

AI Agents Fail 97% of Real Work Tasks: A Human-in-the-Loop Is Still Essential

ended 31. October 2025

Wire reports a new benchmark from Scale AI and the Center for AI Safety just proved what businesses are quietly discovering: AI agents working autonomously are rubbish at real jobs. While basic multi-step automations moving data from one place to another are worthwhile, tasks that need judgment and nuance need humans.

Researchers tested leading AI tools (ChatGPT, Claude, Gemini, and others) on actual freelance work found on sites like Upwork, the kind of project you'd hire a person to do. Tasks included graphic design, video editing, game development, and admin work like data scraping. Each AI got the job brief, the files needed, and an example of what finished work should look like.

The top-performing AI earned just $1,810 out of a possible $143,991 in billable work. That's a 97% failure rate.

Recently, Amazon decided to replace 14,000 jobs with "AI", and Anthropic's CEO predicted 90% coding automation "within months." However, this research reveals that AI struggles when working alone because it can't juggle multiple tools, handle multi-step tasks without drifting, remember sufficient context, or permanently learn from mistakes. It can look busy and churn out some information, but it still can't read the room the same way a human can.

As CAIS director Dan Hendrycks notes, AI models "can't pick up skills on the job like humans" and lack the long-term memory and perceptiveness that real work demands.

This isn't an AI failure. It's proof that humans are wildly optimistic about what AI can achieve. Human judgment, oversight, and adaptation remain essential if endless rework and errors are to be avoided.

The question isn't whether AI can replace workers. It's how we design systems where humans stay in control and AI amplifies what they're already brilliant at doing.

We want your views:

  • Should AI companies come clean and say they're offering plain-old automation which runs on limited logic and basic rules, not AI nuanced decision-making?
  • What tasks have you delegated that a junior human would take in their stride, but turned into an epic struggle with your copilot or automation?
  • Does this 97% failure rate prove that autonomous AI is fundamentally the wrong approach for judgment-based work?
  • If AI needs human oversight for most economically valuable work, how should businesses be restructuring roles rather than eliminating them?
  • What's the dividing line between automation that works (data movement) and automation that fails (complex decision-making)?
  • How do we design transparent review loops that keep humans in control without creating bottlenecks?
  • Where's the accountability when companies blame AI for workforce decisions that still required human strategic choice?
  • Is it fair that individuals find their jobst contorting into QAing a poor AI system, rather then doing the role listed in their job descriptions and contracts.

5 responses from the Newspage community

Copy all

Star Quote
Copy

It's disappointing watching companies rebrand decade-old automation as 'AI automation.' Automation through tools like Zapier and Make has been available for years. Slap on a glossy brochure and hey presto, it's 'innovation'.

Companies are being hoodwinked with snake-oil in their newsfeeds, and frazzled staff are burning a day a week chasing reliability that simply isn't there.

We need to put an end to the wasteful practice of pulling the AI fruit machine lever again and again and again until the answer hits a lacklustre beige jackpot.

Logical automation is excellent. Rekeying data between systems or manually sending standard follow up email sequences? That's madness. Automate it.

But nuanced judgment that replaces human thinking? We're nowhere near AI that's robust and reliable enough yet, and pretending otherwise is disingenuous and burns money and morale.
Copy

Everything and nothing changes. The "paperless office" was heavily forecast from the mid-1970s through the 1990s, with BusinessWeek coining the term around 1975. Into the 1980s, futurists predicted paper would disappear "within a decade." By the 1990s, email, PDFs, and digital document management were supposed to finally eliminate paper "any day now." The reality? Global paper consumption actually increased through the 1990s and early 2000s. The paperless office became a cautionary tale about overestimating technology adoption speed and underestimating human behavior, institutional inertia, and paper's genuine advantages. The 97% AI failure rate isn't surprising—it measures the gap between AI demos and real autonomous work delivery. The core issue: current AI is powerful pattern-matching automation, not an artificial colleague. The path forward isn't elimination but redesign: workflows with humans at judgment nodes and AI handling bounded, repeatable subtasks.
Copy

AI and automation are not the same thing, but too many leaders are being sold as if they are. Automation follows fixed rules. AI learns, reasons and adapts. They work hand in hand, but confusing the two can lead to costly errors. Some marketers deliberately blur the line, slapping an AI label on old automation to cash in on hype. As an AI strategist, I meet numerous business leaders who are dazzled by buzzwords. The real opportunity isn’t replacing people. It’s about redesigning roles so that humans oversee, interpret, and guide what AI does. Research shows AI still fails most “real work” without human judgment. That’s not a flaw. It’s a reminder that oversight is the new superpower. When leaders have the foresight to synthesise AI, automation and humans together, they don’t lose jobs, they reimagine them. Employees gain variety, creativity and control. No two days look the same. The future of work isn’t human or AI. It’s humans and AI, working side by side, smarter together.
Copy

AI agents failing 97% of real jobs shouldn’t shock anyone who’s ever watched one try to send an email without melting down. They’re like interns who can quote Nietzsche but can’t use a printer. The truth is, autonomous AI isn’t replacing humans, it’s just creating more work for us tidying up after it. What it still can’t replicate is the beauty of human emotion: empathy, intuition, humour, and the million invisible choices that come from a life actually lived. That’s the real operating system behind every good decision. The dream of a hands-free office has quickly become a babysitting job for overconfident algorithms, and the only thing they’re automating efficiently is disappointment.
Copy

AI agents can’t work alone and it’s a costly illusion to pretend they can. If 97% of jobs fail when handed to AI we shouldn’t be surprised. The idea that machines can replace judgement, empathy or context is collapsing under its own hype. Automation is fine for shifting data from A to B, but when real-world complexity creeps in, AI falls apart. It can’t read the room, learn from mistakes, or understand consequences. The irony is that many staff are now being paid to quality-check the errors of a system that was supposed to replace them. Let’s stop calling this “AI transformation” and admit what it is: a rule-based shortcut that still depends on human oversight. Until AI can think with, not for, people, it’s creating more expensive time sucks trying to fix its failings.