2026-04-11
GPT-5.4 Is Jagged — And That Kills the General Intelligence Story
GPT-5.4 tops [GPQA Diamond](https://benchmark.space/benchmark/gpqa-diamond) at 92.8%, aces [USAMO 2026](https://benchmark.space/benchmark/usamo-2026) at 95.2%, and scores 0.3% on [ARC-AGI 3](https://benchmark.space/benchmark/arc-agi-3). Same model. Same weights. Same release. The gap between those numbers is the most important fact in AI right now.
The profile that breaks the narrative
OpenAI's own benchmarks tell the story. GPT-5.4 beats the human first attempt on GDPval 70.8% of the time, blind-graded by experts across 44 white-collar occupations. It solved a FrontierMath Tier 4 problem that one mathematician had curated for 20 years — he called it his "personal move 37." On [Mind2Web](https://benchmark.space/benchmark/mind2web) it hits 92.8%. On [BrowseComp](https://benchmark.space/benchmark/browsecomp), 82.7%.
Then the cracks. [SWE-bench Verified](https://benchmark.space/benchmark/swe-bench-verified): 77.2%, behind Claude Mythos Preview at 93.9%. [Humanity's Last Exam](https://benchmark.space/benchmark/humanitys-last-exam): as low as 36.2% depending on the run. [ARC-AGI 3](https://benchmark.space/benchmark/arc-agi-3): 0.3%. That is not a typo. Zero point three percent.
"We are in the spiky world of AI performance where record-breaking performance in one domain derived from the finest of distilled training data does not guarantee that such data exists in another domain." — [AI Explained, 6:16](https://www.youtube.com/watch?v=zizoDORjmlQ&t=376)
The word "spiky" is doing a lot of work. What it really means: there is no general intelligence. There is domain-specific competence that scales with the quantity and quality of available training data for that domain. Code has infinite data and free verification — so models get world-class. Abstract reasoning with novel problem structures has neither — so models flatline.
The old paradigm is dead
Two years ago, the assumption held: if a model improved on one benchmark, it would improve on all of them. Scaling was a rising tide. That assumption is no longer valid.
"In the older paradigm, if a model was clearly better in one domain, it was much more likely that they would be better in many or all domains. That just is not the case anymore." — [AI Explained, 2:02](https://www.youtube.com/watch?v=2_DPnzoiHaY&t=122)
The data confirms it. Look at the [ARC-AGI 2](https://benchmark.space/benchmark/arc-agi-2) leaderboard: [Claude Opus 4.6](https://benchmark.space/model/claude-opus-4.6) leads at 85.0%, GPT-5.4 trails at 73.3%. Flip to [Mind2Web](https://benchmark.space/benchmark/mind2web): GPT-5.4 leads at 92.8%, Claude Opus 4.6 is nowhere near the top. Each model has been optimized for different domains in post-training. The generalist era is over.
| Benchmark | GPT-5.4 | Claude Opus 4.6 | Delta |
| GPQA Diamond | 92.8% | 88.0% | +4.8 |
| ARC-AGI 2 | 73.3% | 85.0% | -11.7 |
| Mind2Web | 92.8% | — | +large |
| SWE-bench Verified | 77.2% | 90.0% | -12.8 |
| USAMO 2026 | 95.2% | 97.6% | -2.4 |
The deltas swing by double digits. This is not a razor-thin frontier where models trade small advantages. This is two different machines built for two different jobs, wearing the same label.
The benchmark gaming multiplier
The jaggedness gets worse when you account for how benchmarks are constructed and scored. Multiple-choice formats inflate scores by 15-20 percentage points compared to open-ended answering with blind grading.
"If you take away the multiple-choice questions, get the models to answer in an open-ended fashion, and then get a blind grader model to compare their answers to the hidden correct answer — you still get some pretty impressive scores, but just not quite as high. Call it a 15 to 20 percentage point drop." — [AI Explained, 8:12](https://www.youtube.com/watch?v=2_DPnzoiHaY&t=492)
That 15-20 point drop is not noise. It means the "general intelligence" you think you're measuring is partly an artifact of the test format. Change the format, change the ranking. Change the domain, change the winner. The model itself hasn't changed — your measurement has.
On ARC-AGI, the problem runs deeper. Models find unintended arithmetic patterns in color-coding grids — shortcuts that produce correct answers without any understanding of the underlying reasoning.
"The numbers representing colors in the input can be used by LLMs to find unintended arithmetic patterns that can lead to accidental correct solutions. I would not call that the models cheating, they are using any shortcuts they can find to get the correct solution." — [AI Explained (citing Melanie Mitchell), 4:04](https://www.youtube.com/watch?v=2_DPnzoiHaY&t=244)
A model that solves ARC-AGI via spurious numeric patterns and a model that solves it via genuine abstract reasoning get the same score. The leaderboard cannot tell them apart.
The Amodei bet
Dario Amodei has staked Anthropic's trajectory on a hypothesis: specialize in enough domains, and generalization emerges. Cover enough of the territory and the gaps fill themselves. It is the most important bet in AI right now, and the GPT-5.4 data is evidence against it.
The FrontierMath Tier 4 result — a 20-year-old problem solved — shows that in domains with deep expert data, models can produce genuinely surprising breakthroughs. But 0.3% on ARC-AGI 3 shows that in domains requiring novel reasoning with thin training data, the same model is helpless. You cannot specialize your way out of a domain where the data does not exist.
"So many of the benchmarks now are written by the labs themselves. They are the only ones with the heft and budget to craft such benchmarks, although obviously that makes them somewhat biased." — [AI Explained, 15:25](https://www.youtube.com/watch?v=2_DPnzoiHaY&t=925)
The labs write the tests, train for the tests, and report the scores. When the model fails a test the lab didn't write — ARC-AGI 3, Humanity's Last Exam — the failure is instructive. The ceiling of what these models can do is not "everything." It is "everything with sufficient training data and verifiable feedback." That is a smaller set than most people assume.
The takeaway
GPT-5.4 is not a general intelligence. It is a collection of domain-specific competencies, some world-class and some sub-human, bundled into a single API endpoint. The same is true of Claude, Gemini, and every other frontier model. The sooner we stop talking about "AGI" as a destination and start talking about domain coverage as a spectrum, the sooner we can have honest conversations about what these systems can and cannot do. The jagged profile is the reality. The smooth curve is the myth.