teortaxesTex on AI benchmarks
12 quotes from AI researchers about benchmarks, models, and evaluation
"Some catching up indeed, Elon. Btw Grok 4.20.690 is still below DeepSeek V3.2-Speciale on CritPt."
Teortaxes @teortaxesTex · 2026-04-08 ·133 likes view on x
"NEW COPE HAS DROPPED. Reuters was right after all... give or take a year. R2/V4/whatever by April 29th?"
Teortaxes @teortaxesTex · 2026-04-10 ·110 likes view on x
"Expert is *exactly* V4-lite. It's faster than V3 has *ever* been, has newer knowledge, and importantly it believes to have 1M context. They are simply using legacy API/V3.2 deployment settings for V4-lite."
Teortaxes @teortaxesTex · 2026-04-08 ·99 likes view on x
"DeepSeek's V4/R2 is one of the most mythologized models ever, probably more than GPT-5; doubly so because there are ZERO leaks, they didn't announce any timeline, any labels or promises. Just noise in an information vacuum."
Teortaxes @teortaxesTex · 2026-04-09 ·87 likes view on x
"DeepSeek has added a whole lot of job listings on Apr 2. Three for agent work, 9 in other R&D, 3 in management, and... 2 in data center operations in Ulanqab City, Inner Mongolia. First open confirmation of new DS-owned compute."
Teortaxes @teortaxesTex · 2026-04-10 ·69 likes view on x
"V4-lite still carries some Speciale sauce in it."
Teortaxes @teortaxesTex · 2026-04-08 ·60 likes view on x
"Qwen *started* this way, Qwen-Plus were closed-first since Qwen 1.5 generation iirc. Kimi, Stepfun were full-closed and only opened after a kick from R1. This is a return to "normalcy"."
Teortaxes @teortaxesTex · 2026-04-09 ·55 likes view on x
"GLM-5 is surprisingly good on prediction arena. Doesnt seem to be mere noise. On the other hand: see GPT 5.4 and Opus 4.6 being well below their official benchmarks — humans can tell."
Teortaxes @teortaxesTex · 2026-04-09 ·45 likes view on x
"Prior to any official announcement from DeepSeek – at least a WeChat post – nothing that's happening is legible. They're brazenly A/B testing, changing inference settings, deploying placeholders in prod."
Teortaxes @teortaxesTex · 2026-04-09 ·45 likes view on x
"DeepSeek has technically stopped open sourcing its frontier models in Q4 2025... so far. Is V4 a fulfillment of Dean's prediction, courtesy of the Party? Or just postponement? Let's wait and see."
Teortaxes @teortaxesTex · 2026-04-09 ·41 likes view on x
"very interesting Only Opus 4.6 and GPT 5.4 manage to absolutely avoid total bankruptcy in long-term betting. This is not so much a question of reasoning as one of learning from mistakes. Clearly, when things start going south, they can adjust towards safety. Open models can't."
"People in policy don't understand the algorithmic improvement rate for post-training in the era of agents with really long contexts. V4 *almost certainly* has the latent capability to do this shit. Maybe with more tokens, maybe even with more total compute."
Teortaxes @teortaxesTex · 2026-04-09 ·32 likes view on x