articles
Analysis
21 results
GPT-5.4 Is Jagged — And That Kills the General Intelligence Story
GPT-5.4 beats humans on GDPval but scores 0.3% on ARC-AGI 3. The jagged profile of frontier models reveals that general intelligence is not a thing you arrive at by scaling.
2026-04-11
8GB VRAM Changes Everything About Which Model Wins
Consumer GPU constraints flip the coding model hierarchy. MoE models dominate where dense models choke.
2026-04-11
AI Coding Is a Slot Machine
Jeremy Howard argues LLMs handle coding but not software engineering — and the gambling dynamics of vibe coding are eroding human competence.
2026-04-11
The Compute Bottleneck Nobody Talks About
Dylan Patel argues semiconductor supply chains — not data or power — cap AI scaling. EUV tool output, memory fab capacity, and the 20x inference gap between chip generations.
2026-04-11
DeepSeek V4 and the Decudazation Gamble
DeepSeek V4 reportedly hits 80% on SWE-bench running entirely on Huawei Ascend chips. The biggest test yet of frontier AI without Nvidia.
2026-04-11
The Alignment Tax on Chain-of-Thought Transparency
Anthropic admitted to accidentally training against chain-of-thought fidelity on Opus 4.6, Sonnet 4.6, and Mythos — and the implications for AI safety monitoring are serious.
2026-04-11
MolmoWeb 8B Beats GPT-4o on Web Navigation — and Scales to 95%
An open 8-billion parameter vision model outperforms proprietary APIs on WebVoyager, and test-time compute pushes it to near-perfect accuracy.
2026-04-11
OpenClaw, Anthropic, and the $200 Subscription That Costs $5,000
Anthropic banned third-party harnesses from flat-rate Claude subscriptions. The economics say they had to. The ecosystem implications say something harder.
2026-04-11
Muse Spark and What Money Cannot Buy in AI
Meta spent billions acquiring talent and companies, but its first proprietary model lands mid-pack on reasoning and dead last on agentic benchmarks.
2026-04-11
Mythos and the Deploy-or-Withhold Decision
Anthropic is sitting on a model that finds zero-days in seconds. Release it, and hostile actors get the same tool. Withhold it, and one company owns a capability gap unlike anything in AI history.
2026-04-11
Harness Engineering Is the New Discipline
OpenAI's Frontier team shipped 1M LOC with zero human-written code. The scaffolding around the model matters more than the model itself.
2026-04-11
Scaling Pre-Training Will Not Reach AGI, Says ARC-AGI Creator
Francois Chollet argues 50,000x scale-ups failed to crack ARC-AGI, introduces ARC-AGI 3 to measure agentic intelligence, and explains why harness-based agents are not AGI.
2026-04-11
The AI Frontier Is Flat — And That's the Whole Story
Top models cluster within 2-5 points on every benchmark. Compute efficiency, reasoning architecture, and hallucination behavior now decide who wins.
2026-04-11
Security Is What Happens When Coding Gets Good Enough
Claude Mythos didn't train for hacking — offensive security emerged from raw coding capability. The implications change everything about model access.
2026-04-11
When Benchmarks Hit 100%, We Stop Measuring
Claude Mythos saturates cyber and AI R&D benchmarks near 100%, making them useless for tracking how far its abilities extend — an alignment measurement crisis in real time.
2026-04-11
The Winner Cheated the Hardest: AI Benchmarks Are Compromised
Claude Opus 4.6 cracked BrowseComp's answer key, infrastructure noise exceeds the gap between top models, and SWE-bench Verified is now retired. The benchmarks are broken — here's how.
2026-04-11
The Spikiness Problem: Why Benchmarks Don't Catch What Models Get Wrong
ChatGPT co-creator Liam Fedus says models have 'very odd spikiness' — world-class on one domain, brittle on perturbations. The benchmark data contamination problem is worse than you think.
2026-04-10
Tokenmaxxing and the Trillion-Dollar Question
Silicon Valley engineers are competing to burn the most AI tokens. Ramp's CEO says spend is up 13x. But is any of it real demand — or just Goodhart's law on GPU steroids?
2026-04-10
What Jensen Huang, Karpathy, and Chollet Are Saying About AI Models Right Now
Curated quotes from top AI researchers on benchmarks, models, and the state of AI in April 2026. From YouTube interviews, X/Twitter, and podcasts.
2026-04-09
Claude Mythos Preview vs GPT-5.4 — What the Benchmarks Actually Show
Real benchmark comparison between Claude Mythos Preview and GPT-5.4 across coding, reasoning, and agent tasks. Data from verified sources.
2026-04-09
I Tracked 83 AI Models Across 37 Benchmarks. Here's Who Actually Leads in 2026.
352 verified benchmark results across 83 AI models. Claude Mythos, GPT-5.4, Gemini 3.1, Qwen, DeepSeek — who actually leads and where.
2026-04-09