- GPT-5 prioritizes smoother user experience over raw capability gains.
- Hallucination rates dropped 45% versus GPT-4o with web search enabled.
- AI benchmarks are approaching saturation, making progress harder to measure.
Sam Altman compared GPT-5 to Apple's Retina display. The analogy was deliberate. When OpenAI released its newest model in August 2025, the company was not claiming a leap in intelligence, according to MIT Technology Review. It was claiming a leap in smoothness.
The distinction matters more than it first appears.
The Model That Decides for You
GPT-5's most notable feature is invisible to the user. An automatic routing system determines whether each query needs fast, direct responses or slower, deeper reasoning, selecting the right approach without asking.
Nick Turley, an OpenAI product lead, acknowledged the shift candidly. GPT-5 "feels better to use," he said, while conceding that good vibes alone would not deliver an automated future. It was the kind of measured claim a product team makes when the technology has outpaced the narrative.
For anyone unfamiliar with the difference between GPT-4o, o1, and o3, the routing system removes a genuine friction point. One model, one interface, no decisions required.
What is model routing?
GPT-5 automatically chooses between fast responses and extended reasoning for each query. Previous OpenAI models required users to select the right model themselves, a distinction most people outside AI circles found baffling.
The practical result is a model that operates faster and costs less than its predecessors. OpenAI offered GPT-5 to non-paying users for the first time, a quiet but notable expansion of access to reasoning-capable AI.
AI Hallucinations Drop, but Benchmarks Tell a Harder Story
The most consequential improvement lies in reliability. With web search enabled, GPT-5 produces roughly 45% fewer factual errors than GPT-4o. In thinking mode, error rates drop further still: approximately 80% fewer hallucinations than OpenAI's o3 model.
Key figure
45%
Reduction in factual errors compared to GPT-4o when GPT-5 uses web search
These numbers carry practical weight. Hallucinated software packages, for instance, could lead users to download malicious code. Reducing confabulation is not a cosmetic upgrade.
Yet the benchmark picture is more sobering. GPT-5 scored 74.9% on SWE-Bench Verified, a coding evaluation. Respectable, but well short of the 80-85% range that would genuinely impress the field.
Clémentine Fourrier, an AI researcher at HuggingFace who studies model evaluation, offered a pointed analogy. Testing frontier models on current benchmarks, she suggested, resembles grading a high schooler on middle-school problems. Success proves little, but failure is revealing.
It's basically like looking at the performance of a high schooler on middle-grade problems. If the high schooler fails, it tells you something, but if it succeeds, it doesn't tell you a lot.
Clémentine Fourrier, HuggingFace
When the Ruler Stops Working
Fourrier's observation points to a deeper problem. The benchmarks themselves are approaching saturation. When multiple frontier models score within a few percentage points of each other, the instruments of measurement fail before the models do.
The pattern recalls an older scientific problem: when thermometers max out, you do not conclude the room has stopped warming. You build a better thermometer.
By March 2026, OpenAI had released GPT-5.2 and GPT-5.4, each bringing incremental gains. GPT-5.4 scored 83% on GDPval, a knowledge work evaluation, and reduced claim-level errors by 33% compared to GPT-5.2.
Key figure
83%
GPT-5.4's score on GDPval knowledge work tasks (March 2026), up from GPT-5.2's 70.9%
Steady refinement, not conceptual rupture.
This trajectory raises an uncomfortable question. If the tools we use to measure AI progress are becoming obsolete, how will we recognize the next genuine advance when it arrives?
More On AI Hallucinations
AI Hallucinations Are Inevitable, OpenAI Researchers Prove
OpenAI researchers prove AI hallucinations are mathematically inevitable. Nine out of ten benchmarks reward guessing over admitting uncertainty.
→OpenAI's answer, for now, appears pragmatic. Rather than chasing a single dramatic benchmark score, the company has shifted toward making existing capabilities more reliable, more accessible, and less prone to the confident fabrications that erode trust.
Several research groups, including Fourrier's team at HuggingFace, are already building harder evaluations to keep pace.
Whether that collective effort will produce clarity or just higher numbers remains to be seen. The old benchmarks, characteristically, have no opinion.
Sources
- Primary Source: GPT-5 is here. Now what? (MIT Technology Review, 2025)
- Additional Context:
- Introducing GPT-5 (OpenAI)
- GPT-5 Benchmarks (Vellum)
- Introducing GPT-5.4 (OpenAI, March 2026)
