AI Agents Fail 97% of Real Work Tasks: A Human-in-the-Loop Is Still Essential
Wire reports a new benchmark from Scale AI and the Center for AI Safety just proved what businesses are quietly discovering: AI agents working autonomously are rubbish at real jobs. While basic multi-step automations moving data from one place to another are worthwhile, tasks that need judgment and nuance need humans.
Researchers tested leading AI tools (ChatGPT, Claude, Gemini, and others) on actual freelance work found on sites like Upwork, the kind of project you'd hire a person to do. Tasks included graphic design, video editing, game development, and admin work like data scraping. Each AI got the job brief, the files needed, and an example of what finished work should look like.
The top-performing AI earned just $1,810 out of a possible $143,991 in billable work. That's a 97% failure rate.
Recently, Amazon decided to replace 14,000 jobs with "AI", and Anthropic's CEO predicted 90% coding automation "within months." However, this research reveals that AI struggles when working alone because it can't juggle multiple tools, handle multi-step tasks without drifting, remember sufficient context, or permanently learn from mistakes. It can look busy and churn out some information, but it still can't read the room the same way a human can.
As CAIS director Dan Hendrycks notes, AI models "can't pick up skills on the job like humans" and lack the long-term memory and perceptiveness that real work demands.
This isn't an AI failure. It's proof that humans are wildly optimistic about what AI can achieve. Human judgment, oversight, and adaptation remain essential if endless rework and errors are to be avoided.
The question isn't whether AI can replace workers. It's how we design systems where humans stay in control and AI amplifies what they're already brilliant at doing.
We want your views:
- Should AI companies come clean and say they're offering plain-old automation which runs on limited logic and basic rules, not AI nuanced decision-making?
- What tasks have you delegated that a junior human would take in their stride, but turned into an epic struggle with your copilot or automation?
- Does this 97% failure rate prove that autonomous AI is fundamentally the wrong approach for judgment-based work?
- If AI needs human oversight for most economically valuable work, how should businesses be restructuring roles rather than eliminating them?
- What's the dividing line between automation that works (data movement) and automation that fails (complex decision-making)?
- How do we design transparent review loops that keep humans in control without creating bottlenecks?
- Where's the accountability when companies blame AI for workforce decisions that still required human strategic choice?
- Is it fair that individuals find their jobst contorting into QAing a poor AI system, rather then doing the role listed in their job descriptions and contracts.




