Copy article

Amazon Abandons AI Usage Leaderboard to Measure Pace of Transformation

ended 29. May 2026

The FT reports Amazon just killed an internal AI leaderboard because employees could game it, and the gaming was costing real money.

Amazon's "Kirorank" dashboard scored developers on their AI tool usage. Staff responded predictably: they assigned AI agents to carry out large and multiple low-value tasks to climb the rankings either to punish the system that threatened to replace them, or to ensure they looked like keen adopters.

Senior VP Dave Treadwell told employees to stop "tokenmaxxing" and warned that the behaviour was inflating computing costs. The leaderboard has been taken offline.

Amazon previously set targets for over 80% of developers to use AI weekly, then built a scoreboard to track it. 

Software engineers have pushed back from the start about using tokens as a quality score, arguing the problem runs deeper than tokenmaxxing. Measuring success by token volume is like measuring a weekly shop by how much you spent rather than whether you came home with what was on the list at a decent price and quality. The rush to adopt AI at whatever the cost has led directly to this outcome.

We want your views:

We'd like your views:

  • If you incentivise volume over value, whose fault is it when staff game the system, the employees threatened with replacement or the target-setters?
  • The shift from flat-fee subscription to consumption-based token pricing directly increases costs when employees overconsume. Who was responsible for not setting the company budget? Do frontier  AI companies have a duty-of-care to monitor rogue spending spikes, like a bank might, to prevent costly mistakes?
  • Amazon now measures “normalised deployments”, evidence of AI producing useful code, instead of raw token use. Is that a genuine fix? In a vibecoding environment, who is reviewing the code for value? A human or a machine?
  • UK firms planning major AI rollouts: if Microsoft, Uber, and Amazon have all hit the same runaway cost issues, what makes any mid-sized company confident it won't?

1 responses from the Newspage community

Copy all

Star Quote
Copy

Engineers flagged token efficiency as the real problem long before any leaderboard went live, in forums and threads that nobody with budget authority was reading. The concerns were specific and technical: consumption-based pricing and large language models do not scale the way enterprise software traditionally does.

Summarising a PDF is a fundamentally different proposition to running thousands of agents across a global engineering team. The costs behave differently. The failure modes behave differently. The value per token behaves differently. But the booster juggernaut rolled over every concern, and anyone who raised them got filed under "resistant to change."

Amazon's 80% weekly usage target tells you everything: leadership measured adoption like a rollout of Office 365, not a consumption model where every query has a variable price tag. Now they've killed the leaderboard, but the target culture that built it is still intact. The engineers weren't the problem. The listening was.