HomeThe New IntelligenceGPT-5 Chose Polish Over Power, and That Tells Us Something

GPT-5 Chose Polish Over Power, and That Tells Us Something

OpenAI's GPT-5 prioritizes user experience over raw power, raising questions about measuring true AI progress.

Share
The New Intelligence · Explore this series
August 8, 2025
Key Takeaways
  • GPT-5 prioritizes smoother user experience over raw capability gains.
  • Hallucination rates dropped 45% versus GPT-4o with web search enabled.
  • AI benchmarks are approaching saturation, making progress harder to measure.

Sam Altman compared GPT-5 to Apple's Retina display. The analogy was deliberate. When OpenAI released its newest model in August 2025, the company was not claiming a leap in intelligence, according to MIT Technology Review. It was claiming a leap in smoothness.

The distinction matters more than it first appears.

The Model That Decides for You

GPT-5's most notable feature is invisible to the user. An automatic routing system determines whether each query needs fast, direct responses or slower, deeper reasoning, selecting the right approach without asking.

Nick Turley, an OpenAI product lead, acknowledged the shift candidly. GPT-5 "feels better to use," he said, while conceding that good vibes alone would not deliver an automated future. It was the kind of measured claim a product team makes when the technology has outpaced the narrative.

For anyone unfamiliar with the difference between GPT-4o, o1, and o3, the routing system removes a genuine friction point. One model, one interface, no decisions required.

What is model routing?

GPT-5 automatically chooses between fast responses and extended reasoning for each query. Previous OpenAI models required users to select the right model themselves, a distinction most people outside AI circles found baffling.

The practical result is a model that operates faster and costs less than its predecessors. OpenAI offered GPT-5 to non-paying users for the first time, a quiet but notable expansion of access to reasoning-capable AI.

AI Hallucinations Drop, but Benchmarks Tell a Harder Story

The most consequential improvement lies in reliability. With web search enabled, GPT-5 produces roughly 45% fewer factual errors than GPT-4o. In thinking mode, error rates drop further still: approximately 80% fewer hallucinations than OpenAI's o3 model.

Key figure

45%

Reduction in factual errors compared to GPT-4o when GPT-5 uses web search

These numbers carry practical weight. Hallucinated software packages, for instance, could lead users to download malicious code. Reducing confabulation is not a cosmetic upgrade.

Yet the benchmark picture is more sobering. GPT-5 scored 74.9% on SWE-Bench Verified, a coding evaluation. Respectable, but well short of the 80-85% range that would genuinely impress the field.

Clémentine Fourrier, an AI researcher at HuggingFace who studies model evaluation, offered a pointed analogy. Testing frontier models on current benchmarks, she suggested, resembles grading a high schooler on middle-school problems. Success proves little, but failure is revealing.

It's basically like looking at the performance of a high schooler on middle-grade problems. If the high schooler fails, it tells you something, but if it succeeds, it doesn't tell you a lot.

Clémentine Fourrier, HuggingFace

When the Ruler Stops Working

Fourrier's observation points to a deeper problem. The benchmarks themselves are approaching saturation. When multiple frontier models score within a few percentage points of each other, the instruments of measurement fail before the models do.

The pattern recalls an older scientific problem: when thermometers max out, you do not conclude the room has stopped warming. You build a better thermometer.

By March 2026, OpenAI had released GPT-5.2 and GPT-5.4, each bringing incremental gains. GPT-5.4 scored 83% on GDPval, a knowledge work evaluation, and reduced claim-level errors by 33% compared to GPT-5.2.

Key figure

83%

GPT-5.4's score on GDPval knowledge work tasks (March 2026), up from GPT-5.2's 70.9%

Steady refinement, not conceptual rupture.

This trajectory raises an uncomfortable question. If the tools we use to measure AI progress are becoming obsolete, how will we recognize the next genuine advance when it arrives?

More On AI Hallucinations

AI Hallucinations Are Inevitable, OpenAI Researchers Prove

OpenAI researchers prove AI hallucinations are mathematically inevitable. Nine out of ten benchmarks reward guessing over admitting uncertainty.

OpenAI's answer, for now, appears pragmatic. Rather than chasing a single dramatic benchmark score, the company has shifted toward making existing capabilities more reliable, more accessible, and less prone to the confident fabrications that erode trust.

Several research groups, including Fourrier's team at HuggingFace, are already building harder evaluations to keep pace.

Whether that collective effort will produce clarity or just higher numbers remains to be seen. The old benchmarks, characteristically, have no opinion.

Sources

Share
Related Articles
Why We Can Never Prove That Someone Else is Conscious

'Rival' scientists use category theory to show that while 'shapes' of experiences might be matched across minds, we can never observe the feeling itself.

AI Consciousness Is Unlikely, Says Neuroscientist Anil Seth

Neuroscientist Anil Seth argues AI consciousness is unlikely without biology. His TED talk lands amid a widening debate over conscious AI, not intuition.

AI In Science Connects the Dots, But Only In Fields That Are Fragmented

An analysis of 80 million papers shows AI boosts originality where knowledge is scattered and connections are weak, but contributes little novelty in structured science.

"Keep Humanity Safe From AI," Urges Pope Leo XIV

Pope Leo XIV's first encyclical reaches the same verdict on AI as the labs building it, then parts ways over the meaning of human limits.