Liam Fedus co-created ChatGPT. He was VP of Post-Training at OpenAI. Now he runs Periodic Labs, applying AI to material science. And he thinks we're fooling ourselves about how capable these models really are.
The most striking observation from Fedus's interview on [No Priors](https://www.youtube.com/watch?v=Oru2Jxr1xHU):
"We've consistently seen these systems have a very odd spikiness and it's actually possible to architect a system that is world-class on some math domain but then you could do some perturbations to the questions and actually degrade it substantially. So it's like a bad high school student." — [Liam Fedus, 22:37](https://www.youtube.com/watch?v=Oru2Jxr1xHU&t=1357)
This is the benchmark contamination problem in a sentence. A model can top the [GPQA Diamond](https://benchmark.space/benchmark/gpqa-diamond) leaderboard and still crumble when you rephrase the same question. The leaderboard shows 94.6% for [Claude Mythos Preview](https://benchmark.space/model/claude-mythos-preview). What it doesn't show is how narrow the competence window actually is.
Fedus revealed an origin story that illustrates how contingent even the biggest AI products are:
"The goal was we need to come up with some productionization of GPT-4. OpenAI had GPT-4, it was pre-trained and there were some rough post-trains on it, and there's questions about how do we turn this incredibly powerful model into products. Some of our least interesting ideas were a meeting bot. But John Schulman was very opinionated. He's like, 'We think we should keep it very general. Let's do a chatbot.'" — [Liam Fedus, 4:06](https://www.youtube.com/watch?v=Oru2Jxr1xHU&t=246)
The most important product in AI history was one opinionated researcher away from being Otter.ai.
Working in material science, Fedus encountered a problem that sounds familiar to anyone following the [benchmark contamination debate](https://benchmark.space/benchmark/mmlu-pro):
"One of the engineers on our team was looking at a reported material property and it was just sort of extracted values from literature and it was really interesting to see the reported value spanned many orders of magnitude. And so you train an ML system on that and the best you can do is model this distribution but you're no closer to a ground truth." — [Liam Fedus, 9:15](https://www.youtube.com/watch?v=Oru2Jxr1xHU&t=555)
This is the scientific version of what happens with MMLU and GSM8K: the training data contains the answers, but the answers themselves are unreliable. MIT Technology Review reported 45% data overlap on QA benchmarks. Fedus is seeing the same problem in a domain where the stakes are physical — wrong material properties mean wrong bridges.
Fedus pushed back on the AGI narrative:
"I think sometimes there's like this mythology of AGI, ASI, RSI and I think we see increasingly powerful systems but they do become limited if they don't have access to the raw data to actually make informed decisions." — [Liam Fedus, 7:11](https://www.youtube.com/watch?v=Oru2Jxr1xHU&t=431)
And offered a more grounded take on what real self-improvement looks like:
"One way I think about recursive self-improvement is really kind of akin to neural architecture search from roughly 10 years ago. And I think there's a very clear path for software engineering. These systems have become so incredibly impressive on this domain as a result of huge amounts of data, really cheap verifiable environments." — [Liam Fedus, 23:39](https://www.youtube.com/watch?v=Oru2Jxr1xHU&t=1419)
The implication: [SWE-bench Verified](https://benchmark.space/benchmark/swe-bench-verified) keeps climbing because code has two things physics doesn't — infinite training data and nearly free verification. Benchmarks in other domains won't follow the same curve.
On whether scaling laws transfer from language to the physical world:
"Everything is about scaling. It's given that predictability. It's allowed us to put huge amounts of capital into this field. And I think the physical sciences, physical engineering will have a very similar property where we establish these scaling properties and bring that mindset." — [Liam Fedus, 18:31](https://www.youtube.com/watch?v=Oru2Jxr1xHU&t=1111)
Fedus is betting his company on this. But he's also clear about the catch: in language, the internet provides the data for free. In material science, every data point requires a physical experiment. The scaling law may hold — but the cost per data point is orders of magnitude higher.
When the person who built ChatGPT says models are "like a bad high school student" under perturbation, that's not FUD — that's an engineering observation from someone who's seen the weights up close. Benchmarks measure the spike, not the valley between spikes. The gap between what [leaderboards](https://benchmark.space/rankings) show and what models actually do reliably may be the most important unmeasured quantity in AI.