3 quotes from AI researchers about benchmarks, models, and evaluation
"With adaptive thinking, maximum effort, and Python tools, Claude Mythos preview scored almost 93% on finding specific UI elements in high-resolution screenshots. That's 10% higher than Claude Opus 4.6."
"To even double AI progress when you factor in compute being a key limiting ingredient, Anthropic predicts that you would actually need an uplift roughly 10 times larger than that 4x one. They think it would take a 40x productivity improvement to see a 2x progress speedup at Anthropic given the crazy bottleneck of compute."