The Reinforcement Gap: Why AI is Mastering Code but Stalling at Poetry
Techcrunch has reported AI’s progress is wildly uneven. Systems like ChatGPT-5, Gemini 2.5 and Sonnet 4.5 can now debug, refactor, and build software faster than most humans. Yet in the creative arts, the story is very different. Ask AI to write a poem or screenplay and you’ll get roughly the same quality of result you did a year ago. One field is accelerating at breakneck speed and the other is stuck in neutral.
The Reinforcement Gap: Measurable vs Meaningful
According to METR (Model Evaluation & Threat Research), the length of software tasks frontier AI agents can complete autonomously with 50 per cent reliability has been doubling every seven months. That exponential progress is driven by reinforcement learning (RL), the same method that a manager uses to train a subordinate, offering feedback and praise for getting closer to the correct answer, again and again.
STEM disciplines are perfect for this. Code either compiles or it doesn’t; a physics simulation obeys gravity or it breaks. Those clear pass–fail conditions create billions of testable outcomes, fuelling continuous, measurable improvement.
Creative fields don’t work that way. Poetry, prose, and screenwriting don’t have unambiguous success metrics. One person’s “padding” is another’s “immersive world-building”. There’s no objective “creativity score” to optimise against. The result is a widening reinforcement gap. The measurable is advancing because it can be measured, while the meaningful is stalling because it can’t.
When Investment Follows the Metrics
Silicon Valley is doubling down on what’s easy to test. Billion-dollar RL “environments” are teaching agents to ace synthetic benchmarks, solve multi-step puzzles and build ever-larger software systems, but not to hold a convincing, compassionate conversation. Funding is flooding into functionality that knows every rule of logic yet still struggles to feel.
Anyone who’s tried to get AI to write humour, irony, or human warmth, or experienced the fight redrafting the company “About Us” page, knows the problem.
The models can pass a coding test but not a Turing-grade moment of connection.
We’d like your views:
- If reinforcement learning only rewards what’s measurable, are we designing an economy that values compliance over creativity?
- Can AI ever master the messy brilliance of poetry, humour, or human judgement?
Should we be developing new, human-centred metrics that recognise depth and meaning, not just mechanical accuracy? - And how do we train the next generation of “AI architects” to preserve the creative spark that code alone can’t generate?
Because if we don’t, we may end up with a world that’s technically brilliant and emotionally tone-deaf.
We design these systems. We own the gap.

