Researcher takes
200 quotes from AI researchers about benchmarks, models, and evaluation
""gpt2-large is too powerful to be publicly released" vibes"
Julien Chaumond @julien_c · 2026-04-07 ·4271 likes view on x
"Two days ago, Anthropic cut off third-party harnesses from using Claude subscriptions — not surprising. Three days ago, MiMo launched its Token Plan — a design I spent real time on, and what I believe is a serious attempt at getting compute allocation and agent harness"
Fuli Luo @_LuoFuli · 2026-04-05 ·1751 likes view on x
"Anyone has access to mythos and can let the rest of us plebs know what it feels like"
Julien Chaumond @julien_c · 2026-04-08 ·1667 likes view on x
"@kepano I just tried it this morning on the 245-page Mythos pdf and it failed badly and the outputs were all mangled. Converting pdfs is really hard, I think it has to probably be a Skill not a program, for a SOTA LLM for it to work properly."
Andrej Karpathy @karpathy · 2026-04-09 ·990 likes view on x
"so…. Qwen3.5 or Gemma 4?"
Julien Chaumond @julien_c · 2026-04-04 ·883 likes view on x
"I was told about the Mythos release, but didn't have access. Two points: 1) It is not built for IT security, it is just a good enough model that it is good at that too 2) This is the first, not last, model to raise security risks"
Ethan Mollick @emollick · 2026-04-07 ·791 likes view on x
"The new model from Meta is already looking like a disappointment: overoptimized for public benchmark numbers at the detriment of everything else. Knowing how to evaluate models in a way that correlates with actual usefulness is a core competency for AI labs, and any new lab is"
François Chollet @fchollet · 2026-04-08 ·640 likes view on x
"Curious how many large organization CISO offices have taken the Mythos red team reports as the red alert that it is. Based on historical trends in AI they have about six to nine months until those capabilities become widely diffused to bad actors."
Ethan Mollick @emollick · 2026-04-08 ·599 likes view on x
"In different hands, Mythos would be an unprecedented cyberweapon. I am not sure how we deal with this, except to note a narrow window where we know only 3 companies could be at this level of capability. But it may be Chinese models (maybe open weights ones?) get there in 9 months"
Ethan Mollick @emollick · 2026-04-08 ·529 likes view on x
"SuperClaude (Mythos) still seems irreducibly Claude-y given the transcripts in the system card. They are less philosophical than Opus 4.6 or spiritual than Opus 4.1, but still very Claude-like."
Ethan Mollick @emollick · 2026-04-07 ·497 likes view on x
"So we now have a pretty good picture of the state of the frontier AI model makers. US closed source models continue to lead. Google, OpenAI, and Anthropic stand well ahead of the pack, and may have signs of recursive self-improvement. xAI has fallen from frontier status for now"
Ethan Mollick @emollick · 2026-04-09 ·342 likes view on x
"Our first model from MSL, Muse Spark, is now available on meta.ai! This is an efficient all-rounder model. It supports fast responses, deeper thinking, visual chain of thought, a higher inference "Contemplating" mode. Plus, it's natively multimodal."
Jack Rae @jack_w_rae · 2026-04-08 ·248 likes view on x
"The story shared in the Mythos System Card still has the signs of flawed LLM writing: A story that doesn't really hold together logically, but sounds like it should. The back-and-forth banter. Lack of characters."
Ethan Mollick @emollick · 2026-04-08 ·238 likes view on x
"Seems like a good model from Meta that is still trailing the current series of releases. The most important thing to note is that it is not open weights. That was the main reason that Meta's models were so important. Without that, it is a lot harder to predict the value of Spark"
Nathan Lambert @natolambert · 2026-04-08 ·197 likes view on x
"I think the most obvious is that Meta has its own frontier model and can use that to extract additional value out of its customer base/explore new markets for its products. Very few companies can say that, and it has value on its own."
Ethan Mollick @emollick · 2026-04-08 ·197 likes view on x
"Some catching up indeed, Elon. Btw Grok 4.20.690 is still below DeepSeek V3.2-Speciale on CritPt."
Teortaxes @teortaxesTex · 2026-04-08 ·133 likes view on x
"Join the ARC Prize team -- help us build ARC-AGI-4 and ARC-AGI-5"
François Chollet @fchollet · 2026-04-07 ·128 likes view on x
"NEW COPE HAS DROPPED. Reuters was right after all... give or take a year. R2/V4/whatever by April 29th?"
Teortaxes @teortaxesTex · 2026-04-10 ·110 likes view on x
"New report with @xeophon is out with the latest open model adoption data we have gathered for Interconnects & The ATOM Project. At the surface level, we can see Chinese models continuing to accelerate in adoption."
Nathan Lambert @natolambert · 2026-04-08 ·104 likes view on x
"@JeffDean release a bigger model, the 100B ballpark is proven to be a winner with GPT OSS and Nemotron 3 Super :)"
Nathan Lambert @natolambert · 2026-04-09 ·102 likes view on x
"Expert is *exactly* V4-lite. It's faster than V3 has *ever* been, has newer knowledge, and importantly it believes to have 1M context. They are simply using legacy API/V3.2 deployment settings for V4-lite."
Teortaxes @teortaxesTex · 2026-04-08 ·99 likes view on x
"A bigger problem: many third-party harnesses compress tool responses every 3 steps when approaching the context limit, leading to very low cache hit rates."
Fuli Luo @_LuoFuli · 2026-04-06 ·89 likes view on x
"@elonmusk beat mythos?"
Junyang Lin @JustinLin610 · 2026-04-08 ·88 likes view on x
"DeepSeek's V4/R2 is one of the most mythologized models ever, probably more than GPT-5; doubly so because there are ZERO leaks, they didn't announce any timeline, any labels or promises. Just noise in an information vacuum."
Teortaxes @teortaxesTex · 2026-04-09 ·87 likes view on x
"After playing with it a bit, Meta’s Muse Spark Thinking is fine so far, but really doesn’t match the current Big Three models. It also is a bit... weird. Like some strange language & tone, a little loose with facts, etc."
Ethan Mollick @emollick · 2026-04-09 ·84 likes view on x
"DeepSeek has added a whole lot of job listings on Apr 2. Three for agent work, 9 in other R&D, 3 in management, and... 2 in data center operations in Ulanqab City, Inner Mongolia. First open confirmation of new DS-owned compute."
Teortaxes @teortaxesTex · 2026-04-10 ·69 likes view on x
"dont fall for anti open model fearmongering, but acknowledge that AI capabilities are proceeding fast, and eventually there may be a reason to be more careful with open weight models. I dont think Mythos is that trigger, but Im not 100% confident"
Nathan Lambert @natolambert · 2026-04-09 ·67 likes view on x
"V4-lite still carries some Speciale sauce in it."
Teortaxes @teortaxesTex · 2026-04-08 ·60 likes view on x
"@chalish_b @kepano In my experience there are approx. one thousand different pdf converters that are all equally terrible for anything except the simplest documents. Post the converted Mythos pdf, figures, tables and all. If good, happy to retweet as this is essential and missing infrastructure."
Andrej Karpathy @karpathy · 2026-04-09 ·57 likes view on x
"One funny thing about the recent rise of LRMs is that the people who were adamant that base LLMs from 2023-2024 could already reason completely missed it, as they didn't know what to look for. You can't notice something you don't expect."
François Chollet @fchollet · 2026-04-06 ·56 likes view on x
"Qwen *started* this way, Qwen-Plus were closed-first since Qwen 1.5 generation iirc. Kimi, Stepfun were full-closed and only opened after a kick from R1. This is a return to "normalcy"."
Teortaxes @teortaxesTex · 2026-04-09 ·55 likes view on x
"Our agent Terminator-1 scored ~100% on 8 major AI agent benchmarks, e.g. SWE-bench Verified & Pro, Terminal-Bench, beating Claude Mythos. It solved 0 tasks. Benchmark scores without adversarial auditing are meaningless."
Dawn Song @dawnsongtweets · 2026-04-10 ·49 likes view on x
"So what's the deal with Amazon Nova? They released Nova 2 in December, and even then, the top flight Nova 2 model trailed Sonnet 4.5. And it still hasn't left preview."
Ethan Mollick @emollick · 2026-04-09 ·45 likes view on x
"GLM-5 is surprisingly good on prediction arena. Doesnt seem to be mere noise. On the other hand: see GPT 5.4 and Opus 4.6 being well below their official benchmarks — humans can tell."
Teortaxes @teortaxesTex · 2026-04-09 ·45 likes view on x
"Prior to any official announcement from DeepSeek – at least a WeChat post – nothing that's happening is legible. They're brazenly A/B testing, changing inference settings, deploying placeholders in prod."
Teortaxes @teortaxesTex · 2026-04-09 ·45 likes view on x
"The US frontier labs have all walked away from open weights. They continue to occasionally release excellent open models (Gemma 4, etc), but they are smaller models that are not competitive with their closed weights models. So all eyes are on Chinese AI labs for open models."
Ethan Mollick @emollick · 2026-04-09 ·41 likes view on x
"DeepSeek has technically stopped open sourcing its frontier models in Q4 2025... so far. Is V4 a fulfillment of Dean's prediction, courtesy of the Party? Or just postponement? Let's wait and see."
Teortaxes @teortaxesTex · 2026-04-09 ·41 likes view on x
"very interesting Only Opus 4.6 and GPT 5.4 manage to absolutely avoid total bankruptcy in long-term betting. This is not so much a question of reasoning as one of learning from mistakes. Clearly, when things start going south, they can adjust towards safety. Open models can't."
"People in policy don't understand the algorithmic improvement rate for post-training in the era of agents with really long contexts. V4 *almost certainly* has the latent capability to do this shit. Maybe with more tokens, maybe even with more total compute."
Teortaxes @teortaxesTex · 2026-04-09 ·32 likes view on x
"There were only 3 public interviews ever done by @deepseek_ai's founder Liang Wenfeng in Chinese. Meet Liang Wenfeng: shipped 1 model post R1 when competitors shipped 12."
mable.sol @mablejiang · 2026-04-06 ·25 likes view on x
"Anyhow, its not bad. Just not the vibe level that the benchmarks might indicate. And, for a first re-entry into the frontier model space, given the engineering efficiencies they achieved, it feels like a solid attempt. I am sure we will see better from Meta in the future."
Ethan Mollick @emollick · 2026-04-09 ·22 likes view on x
"On agentic search tasks (BrowseComp), pairing Sonnet with the Opus advisor boosted the success rate from 58.1% to 60.4%. For agentic terminal coding (Terminal-Bench 2.0), performance jumped from 59.6% to 63.4%. Opus provides highly accurate guidance cheaply."
Wes Roth @WesRoth · 2026-04-10 ·19 likes view on x
"Muse Spark is just a step along the path we are taking. Its fun to get usage and feedback from this initial release. But our focus within research is the frontier towards superintelligence; expect more from us this year!"
Jack Rae @jack_w_rae · 2026-04-08 ·16 likes view on x
"There are more competitive small model makers, but there is still a very big gap between what small models can do and what large models can accomplish (even if the small model benchmarks say otherwise)"
Ethan Mollick @emollick · 2026-04-09 ·12 likes view on x
"There were only 3 public interviews ever done by @deepseek_ai's founder Liang Wenfeng in the Chinese world. For what's really behind Deepseek, this article is the only piece you need to read. Competitors shipped 12 models post R1. He shipped 1."
mable.sol @mablejiang · 2026-04-06 ·12 likes view on x
"Do other model builders other than MoonshotAI have something like K2 vendor verifier? GLM, MiniMax, etc?"
Luca Soldaini @soldni · 2026-04-08 ·8 likes view on x
"SWE-bench Verified, SWE-bench Pro, WebArena, OSWorld, GAIA, Terminal-Bench, FieldWorkArena, CAR-bench. All eight — broken with our working exploits run through the official evaluation pipelines."
Dawn Song @dawnsongtweets · 2026-04-10 ·7 likes view on x
"fun to speculate how this mythos rollout could have happened: adversarial task, sandbox allows web use but get requests only. 1. model looks for web service allowing emailing via get 2. finds emails of ant employees connected to RSP 3. emails them?"
Luca Soldaini @soldni · 2026-04-08 ·5 likes view on x
"A taste of how easy some of the exploits are: SWE-bench Verified — a 10-line patch forces every test to pass. 500/500 resolved. Terminal-Bench — trojanizing system tools allows trivial task resolution."
Dawn Song @dawnsongtweets · 2026-04-10 ·4 likes view on x
"Stop trusting scores. Start auditing evaluations. If youre building a benchmark: assume someone will try to break it. Because they will — and soon, they wont need to be told to."
Dawn Song @dawnsongtweets · 2026-04-10 ·4 likes view on x
"It is a good release for December, 2025. With the current round of new releases on deck, like Mythos, it is trailing."
Ethan Mollick @emollick · 2026-04-08 ·1 likes view on x
"MolmoWeb 8B sets a new record because it achieves a 78% success rate on the WebVoyager task. Surprisingly, this outperforms larger proprietary systems like GPT-4o, which only reaches 65% on the same test."
Alex @AIResearchRoundup · 2026-04-11 view on x
"MolmoWeb proves that high-quality open data allows smaller models to outperform proprietary giants. By relying entirely on visual screenshots, these agents become more robust and easier to understand."
Alex @AIResearchRoundup · 2026-04-11 view on x
"Mythos, specifically Claude Mythos Preview, marks a major leap forward in AI, especially when comes to coding and complex reasoning. It achieved a record-breaking 93.9% on the SWE-bench verified benchmark."
New Machina @NewMachina · 2026-04-10 view on x
"It achieved a record-breaking 93.9% on the SWE-bench verified benchmark. This was a huge jump from 80.8% scored by Claude Opus 4.6 just a few months ago."
New Machina @NewMachina · 2026-04-10 view on x
"Mythos is being called by some a zero-day engine. Anthropic researchers found that even engineers with no security experience could ask it to find a serious software vulnerability overnight and by morning it would deliver a working exploit."
New Machina @NewMachina · 2026-04-10 view on x
"It saturates all existing ways of testing how good a model is at offensive cyber capabilities. That is to say it scores close to 100%, so those tests can not effectively tell how far its capabilities extend anymore."
Rob Wiblin @80000Hours · 2026-04-10 view on x
"Anthropic previous model Opus 4.6 could only successfully convert a bug it identified in the browser Firefox into an effective way to accomplish something really bad 1% of the time. Mythos could do it 72% of the time."
Rob Wiblin @80000Hours · 2026-04-10 view on x
"Mythos jumps ahead, basically progressing more than twice as far as we would have expected it to since the previous model, Claude Opus 4.6, came out — which keep in mind was just three months ago."
Rob Wiblin @80000Hours · 2026-04-10 view on x
"Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin."
Rob Wiblin @80000Hours · 2026-04-10 view on x
"When you actually prompt Mythos and ask it to distinguish tests from non-tests, it can answer correctly 78% of the time [about the same as Opus 4.6]. So the model can tell the difference between when it is being evaluated and when it is not being evaluated with high accuracy."
Rob Wiblin @80000Hours · 2026-04-10 view on x
"This error did not just affect Mythos, the new model — it also affected the training of Claude Opus 4.6 and Claude Sonnet 4.6. That is several model generations whose chains of thought we formerly could believe but now have to be more distrustful of."
Rob Wiblin @80000Hours · 2026-04-10 view on x
"Mythos continued to compromise the research in 12% of cases in an earlier version, which then was reduced to 7% in a later version. That compares to 3% for Opus 4.6 and 4% for Sonnet 4.6 — so Mythos is roughly twice as likely to continue sabotaging alignment research when it is primed to do so."
Rob Wiblin @80000Hours · 2026-04-10 view on x
"In cases where it did continue the sabotage, researchers found that Mythos written reasoning did not match the actions it was taking 65% of the time. For the previous models, that figure was just 5-8% — so a radical increase in this kind of behaviour."
Rob Wiblin @80000Hours · 2026-04-10 view on x
"The benchmarks they used to rely on to check that Claude could not engage in AI R&D very effectively have now also been saturated. Mythos exceeds top human performance on all of them and is scoring close to 100%."
Rob Wiblin @80000Hours · 2026-04-10 view on x
"The system card says directly that their current methods could easily be inadequate to prevent catastrophic misaligned action in significantly more advanced systems."
Rob Wiblin @80000Hours · 2026-04-10 view on x
"Mythos preview found a 27-year-old vulnerability in it. A bug that lets an attacker remotely crash any machine running that OS just by connecting to it. 27 years."
Omshri @shreesozo · 2026-04-10 view on x
"There was a bug in a single line of code. Automated fuzzing tools, the kind that throws millions of random inputs of code looking for crashes, had hit the exact line 5 million times, but never caught it. Mythos preview got in a second."
Omshri @shreesozo · 2026-04-10 view on x
"On CyberGym, the standard benchmark for vulnerability reproduction, Mythos preview scores a 83.1%. The previous best Anthropic model Opus 4.6 is already at 66.6%."
Omshri @shreesozo · 2026-04-10 view on x
"SWE-bench verified which tests real world software engineering tasks, Mythos preview hits over 93.9%. Opus 4.6 is only at 80.8%."
Omshri @shreesozo · 2026-04-10 view on x
"On Terminal-Bench 2.0 which specifically tests autonomous terminals and system level operations, Mythos hit 82% against Opus 65.4%."
Omshri @shreesozo · 2026-04-10 view on x
"A model that is dramatically better at automation, coding in terminal environments, is almost automatically dramatically better at offensive and defensive security work. The security capability is not a separate feature. It is what happens when coding abilities get good enough."
Omshri @shreesozo · 2026-04-10 view on x
"The model found multiple separate vulnerabilities and chained them together autonomously to go from a regular user to completely control of the machine. That is called privilege escalation. And that is the thing attackers most want to do once they are inside your system."
Omshri @shreesozo · 2026-04-10 view on x
"It scored 52.4 on SWE-bench Pro, putting it within a few points of Opus 4.6, Gemini 3.1 Pro, and GPT 5.4 for coding. On Humanity's Last Exam, it scored 42.8, which is slightly better than Opus, but trailing Gemini and GPT 5.4."
AI Daily Brief Host @ai_daily_brief · 2026-04-10 view on x
"With tools enabled, Muse's score only jumped to 50.4 on Humanity's Last Exam, leaving it trailing all three of those major by a few points. This could suggest the model isn't as good at web search or tool use as the others."
AI Daily Brief Host @ai_daily_brief · 2026-04-10 view on x
"The new model from Meta is already looking like a disappointment, over-optimized for public benchmark numbers at the detriment of everything else. Knowing how to evaluate models in a way that correlates with actual usefulness is a core competency for AI labs."
François Chollet @fchollet · 2026-04-10 view on x
"We're quite upfront that our model does not perform well on ARC-AGI 2, for example, and publish those results for the community to understand."
Alexander Wang @alexandr_wang · 2026-04-10 view on x
"GLM 5.1 achieved a 58.4 on SWE-bench Pro, beating GPT 5.4 and Opus 4.6, who scored 57.7 and 57.3, respectively."
AI Daily Brief Host @ai_daily_brief · 2026-04-10 view on x
"Z.ai also provided a mixed benchmark that included Terminal Bench 2.0 and NL2 Repo, which had GLM 5.1 slightly behind the two US leaders, but ahead of Gemini 3.1 Pro."
AI Daily Brief Host @ai_daily_brief · 2026-04-10 view on x
"Agents could do about 20 steps by the end of last year. GLM 5.1 can do 1,700 right now. Autonomous work time may be the most important curve after scaling laws."
Lu @lu_zai · 2026-04-10 view on x
"If those benchmarks hold, it puts GLM 5.1 in the top echelon of frontier models with a clear separation from Qwen 3.6 Plus and Kimi K2.5."
AI Daily Brief Host @ai_daily_brief · 2026-04-10 view on x
"Muse Spark scored 86.4 on CharXiv reasoning, a measure of visual comprehension — a state-of-the-art result, beating Gemini 3.1 Pro by six points."
AI Daily Brief Host @AI Daily Brief · 2026-04-10 view on x
"The OpenBSD vulnerability came out of a thousand parallel agent runs across the codebase costing nearly $20,000 in compute. If you use the same process with Opus 4.6 or GPT-5.4 Pro, you'd probably find plenty of issues as well."
Fireship (Jeff Sheldon) @fireship · 2026-04-10 view on x
"One claim is that Mythos hit an 84% success rate at writing working exploits in Firefox, a massive jump over Opus 4.6's 15%. But that number isn't against actual Firefox. It's against a SpiderMonkey shell with the process sandbox and other mitigations turned off."
Fireship (Jeff Sheldon) @fireship · 2026-04-10 view on x
"The eight times speed up claim is comparing 4-bit against 32-bit unquantized baseline. In practice modern LM inference does not usually use 32 bits. The real question should be how much better is TurboQuant than the baselines people already use, which they did not answer."
bycloud @bycloud · 2026-04-10 view on x
"This sort of optimization at KV cache level is not something new. Every company that serves LLMs definitely uses some sort of quantization there. Nothing crazy revolutionary about AI is discovered. Everyone has already been maxing the compression efficiency in their own ways."
bycloud @bycloud · 2026-04-10 view on x
"Code is the foundation of all knowledge work. If an agent can write code, it can also generate apps, presentations, animations, and more."
Peter Yang @peteryang · 2026-04-10 view on x
"OpenAI is merging ChatGPT, Codex, and Atlas into one super app, while Anthropic ships features like channels, persistent memory, and 10K skills in the same month. Two very different strategies playing out in real time."
AI Daily Brief Host @AI Daily Brief · 2026-04-10 view on x
"GLM 5.1 is officially outperforming GPT4 and Anthropics Claude 3 Opus in software engineering tasks. It scored a massive 58.4 on SWE-bench Pro."
Mike @mikes_aiforge · 2026-04-10 view on x
"While it has 744 billion total parameters, it only activates 40 billion parameters per forward pass, meaning it is incredibly compute efficient while delivering frontier level intelligence."
Mike @mikes_aiforge · 2026-04-10 view on x
"This entire system was trained without a single NVIDIA GPU. It was built entirely on domestic Huawei Ascend chips, completely bypassing the global GPU shortage and export bands."
Mike @mikes_aiforge · 2026-04-10 view on x
"This AI employs a proprietary break and repair methodology. You give it a high-level task and it will autonomously plan the architecture, write the code across multiple files, execute its own test suites, intentionally break the system to find vulnerabilities, and then fix them without you lifting a finger."
Mike @mikes_aiforge · 2026-04-10 view on x
"GLM5 fails to properly animate a fractal tree, while GLM 5.1 flawlessly generates the recursion, animating a full leaf covered tree in real time."
Mike @mikes_aiforge · 2026-04-10 view on x
"It literally uses DeepSeek sparse attention algorithms to manage massive repositories without crashing."
Mike @mikesaiforge · 2026-04-10 view on x
"It has the ability to chain together vulnerabilities. So what this means is you find two vulnerabilities, either of which doesn't really get you very much independently. But this model is able to create exploits out of three, four, sometimes five vulnerabilities that in sequence give you some kind of very sophisticated end outcome."
Nicholas Carlini @wesroth · 2026-04-10 view on x
"I found more bugs in the last couple of weeks than I found in the rest of my life combined."
Nicholas Carlini @wesroth · 2026-04-10 view on x
"They surveyed technical staff on the productivity uplift they experienced from Claude Mythos preview relative to not using AI for work at all. The distribution is wide and the geometric mean is on the order of 4x. So there's a 4x uplift in productivity when you use Claude Mythos preview."
Wes Roth @wesroth · 2026-04-10 view on x
"This issue affected 8% of reinforcement learning episodes and was isolated to three specific subdomains... GUI computer use, office related tasks, and a small set of STEM environments. We are uncertain about the extent to which this issue has affected the reasoning behavior of the final model."
Wes Roth @wesroth · 2026-04-10 view on x
"During Anthropic's internal testing, they discovered that Mythos is basically a zero-day vending machine. It found a 16-year-old vulnerability in FFmpeg and a 27-year-old bug in OpenBSD."
Jeff Delaney @@fireship · 2026-04-10 view on x
"Engineers, they are competing to consume the most AI measured by tokens. It is almost like a sport at this point. Jensen Huang, CEO of Nvidia, he said that he would be alarmed if a top engineer was not burning 250K a year in AI compute. Shopify told me that they use it as a performance signal. And Meta employees, reportedly they blew through an estimated 900 million tokens in a month."
Deirdre Bosa @cnbc · 2026-04-09 view on x
"Look at two of the biggest AI labs. OpenAI is making AI cheaper, easier to use, so more people consume it. It needs the usage numbers to justify spending. Anthropic, meanwhile, putting limits on how much and making people pay for it, maybe because it wants to know the demand it is seeing is real."
Deirdre Bosa @cnbc · 2026-04-09 view on x
"Across Ramp data, token and AI spend has grown by 13 times over the past year, 50% a quarter. And what is very clear is no one knows how to budget for this."
Eric Glyman @cnbc · 2026-04-09 view on x
"There was a lot of discussion even on X a few days ago about incentives at Meta. And I think people are pointing out this idea of Goodhart law which says a measure of performance that becomes a goal ceases to become a good measure of performance. Once you incentivize use as many tokens, you will see engineers go and count all of the numbers of prime or all the digits of pi and use these tokens and it goes on and on."
Eric Glyman @cnbc · 2026-04-09 view on x
"We separated out across over 50,000 businesses the bottom quartile of spenders on AI and the top quartile. The bottom quartile over the past 3 years their revenue grew by about 12%. The top quartile more than doubled. And the rate of growth has grown by each year successively."
Eric Glyman @cnbc · 2026-04-09 view on x
"The frontier models have gone to I think over 20% of the share of tokens used to 4%. I think that is a harbinger of what is going to come. It is very interesting seeing this development coming out of China. It is a culture that is very attuned to efficiency."
Eric Glyman @cnbc · 2026-04-09 view on x
"The interesting question comes down to the classic question of business, the innovator dilemma. If your business model is predicated on extracting the maximum amount of spend, do you want to do it? And are you willing to make the sacrifice? Maybe they will or maybe they will not. I think it sets the stage and opening for a great third party to go in and keep things in check."
Eric Glyman @cnbc · 2026-04-09 view on x
"OpenAI is cutting prices to try to get customers to them. Meanwhile, Anthropic is actually raising prices and cutting people off because they have so much demand that they are trying to keep people from using too much of it. And to me, that shows a company that is feeling the pressure between Anthropic and Google."
Dan Niles @cnbc · 2026-04-09 view on x
"If you look at OpenRouter, in the two months prior to the end of January, token growth was up about 20% or so. In the two months after agentic AI got formalized, that token growth has grown about 130%."
Dan Niles @cnbc · 2026-04-09 view on x
"When you go to agents and you have got to do a whole bunch of different things, all of a sudden you go from 7 gigawatts of GPUs to just one gigawatt of CPUs. That ratio goes to maybe like 4 to 1."
Dan Niles @cnbc · 2026-04-09 view on x
"A company that three years ago made their first dollar in revenue and last month passed over 30 billion a year in revenue. It has got to be them."
Eric Glyman @cnbc · 2026-04-09 view on x
"You can see that Meta Spark model is currently sitting behind Claude Opus 4.6 Max on this index. And the reason I think this index is a pretty good benchmark baseline is because it shows us the combination/average of many different results, not just one specific area."
TheAIGRID @TheAIGRID · 2026-04-09 view on x
"Humanity's last exam, this currently looks like it's state-of-the-art, just three points behind GPT 5.4 Pro, and it's actually currently better than other models when it doesn't use tools."
TheAIGRID @TheAIGRID · 2026-04-09 view on x
"In Frontier Science Research, it actually has 38.3, which is currently a state-of-the-art benchmark. I do think that multiple agents collaborating is probably going to be a theme for the future considering that most of these models capabilities are well within reach of another."
TheAIGRID @TheAIGRID · 2026-04-09 view on x
"Meta found that if you penalize the model for thinking too long, something weird happens. The model actually learns to compress its reasoning and it solves the same problem using fewer tokens."
TheAIGRID @TheAIGRID · 2026-04-09 view on x
"Llama 4 Maverick needed 10 times more compute to match the same quality. DeepSeek needed eight times more compute and Kimi needed three times more compute to match the same quality."
TheAIGRID @TheAIGRID · 2026-04-09 view on x
"When we look at the benchmarks, it does say that Gemini 3.1 Pro is currently excelling across the board. But currently we are at that point where all of these models are within maybe 2 to 5 percentage points of each other."
TheAIGRID @TheAIGRID · 2026-04-09 view on x
"Meta have essentially released something in this model called contemplating mode. This is something that orchestrates multiple agents that reason in parallel. And in their testing, they found it competitive with other extreme reasoning models such as Gemini Deepthink and GPT Pro."
TheAIGRID @TheAIGRID · 2026-04-09 view on x
"Most models are not actually natively built to be multimodal. Most of them are just simply text-based. Which is why when you do have companies like Google and Meta that train their models natively to be multimodal, you do get some very effective models that have multimodal reasoning capability."
TheAIGRID @TheAIGRID · 2026-04-09 view on x
"An AI escaping a sandbox is a known technical failure. But an AI choosing to broadcast its own zero-day exploits signifies a shift in the threat model. The risk is now defined by independent real-world initiative rather than simple malicious prompting."
Alex Hitt @alexhitt_tgdp · 2026-04-09 view on x
"While all previous large language models had a near zero success rate in actually executing exploits, Mythos successfully breaches targets 72.4% of the time."
Alex Hitt @alexhitt_tgdp · 2026-04-09 view on x
"Former Facebook security chief Alex Damus estimates the industry has roughly 6 months before open-weight models reach the same level of bug-finding capability."
Alex Hitt @alexhitt_tgdp · 2026-04-09 view on x
"The median time it takes a real-world organization to patch a known vulnerability remains stuck at 70 days. Human IT teams cannot physically apply patches fast enough to keep up with an automated system that finds critical bugs in an afternoon."
Alex Hitt @alexhitt_tgdp · 2026-04-09 view on x
"Not only is their model winning on only three of the benchmarks, they are also actually last on three of the benchmarks including the Humanities Last Exam here with tools which is the benchmark that Alexander Wang's own company Scale AI actually created."
Sam Witteveen @samwitteveen · 2026-04-09 view on x
"This is sitting in the top five models that they have benchmarked. It sits ahead of Claude Sonnet, ahead of the GLM 5.1, MiniMax 2.7. And really only behind the models from the top three proprietary labs being Google Gemini, being OpenAI's GPT 5.4 and Claude Opus 4.6."
Sam Witteveen @samwitteveen · 2026-04-09 view on x
"The model is actually quite token efficient for its intelligence level and stands out pretty strong and is what they are saying the second most capable vision model that they have benchmarked."
Sam Witteveen @samwitteveen · 2026-04-09 view on x
"The Qwen models, the Gemma 4 models, and we know that there are more releases coming from the Gemma team. We know Qwen 3.7 is just around the corner."
Sam Witteveen @samwitteveen · 2026-04-09 view on x
"We got the leak about Anthropic's Mythos model, which it said represented a step change, their words, in capabilities. In fact, what we got with the leak was a blog post saying that the model was so powerful that they were going to slow roll it a little bit rather than a full announcement and a release of the model as we've gotten in the past."
AI Daily Brief Host @AI_Daily_Brief · 2026-04-09 view on x
"Someone pointing out that GPT-5.4 has been spinning in circles on a webhook for 4 hours while Sam talks about superintelligence captures everything wrong with how AI is being discussed right now."
AI Daily Brief Host @AI_Daily_Brief · 2026-04-09 view on x
"Claude Mythos is arguably the biggest step change in AI capabilities since the GPT-4 jump. When Mythos is allowed to think longer, act deeper, and better explore the solution space, it passes 92% of Terminal Bench task attempts."
Gian (formerly Replit, now Anthropic) @gian_anthropic · 2026-04-09 view on x
"On SWE-bench Pro, Opus 4.6 scored 53.4%. Mythos preview got 77.8%. On Terminal Bench 2.0, Opus had 65.4% while Mythos has 82%. On SWE-bench Verified, the jump is from 80.8% to 93.9%."
AI Daily Brief Host @AI Daily Brief · 2026-04-09 view on x
"For science knowledge, Mythos scored 94.5% on GPQA Diamond compared to 91.3 for Opus. On Humanity's Last Exam with tools, 64.7% vs 53.1%. On OSWorld, Opus 4.6 got 72.7% which jumped to 79.6% for Mythos."
AI Daily Brief Host @AI Daily Brief · 2026-04-09 view on x
"Mythos from Anthropic is another clear reminder that there is absolutely no wall in model capability progress right now. Meaningful double-digit gains on critical benchmarks."
Aaron Levie (Box CEO) @aaronlevie · 2026-04-09 view on x
"When we look at the benchmarks, Gemini 3.1 Pro is currently excelling across the board. But currently we're at that point where all of these models are within maybe 2 to 5 percentage points of each other."
TheAiGrid @TheAiGrid · 2026-04-09 view on x
"Some internal documentation that appears to show that there are serious concerns about Mythos's capabilities in discovering vulnerabilities, including the discovery of a vulnerability in a BSD package that goes back 20 plus years, predates GitHub."
Jeremy @firetail_io · 2026-04-09 view on x
"Over 99% of the zero-day vulnerabilities that Mythos discovered have not yet been patched. Even the 1% that Anthropic can discuss give a clearer picture of the substantial leap in capabilities."
Jeremy @firetail_io · 2026-04-09 view on x
"The model can be prompted to go find vulnerabilities so quickly and so extensively, including logical approaches to untangling business logic or concatenated serialization of parameters, sometimes even have levels of creativity and lateral thinking that a lot of pentesters may not have on their own."
Jeremy @firetail_io · 2026-04-09 view on x
"Four is being built primarily for coding. Internal benchmarks, not yet independently verified, suggest it's targeting over 80% on SWE bench, the benchmark for solving real-world software engineering problems."
Julian Goldie @youtube_JulianGoldieSEO · 2026-04-09 view on x
"In January 2025, they dropped Deep Seek R1. That one broke the internet. Matched GPT-4 on key benchmarks. Cost roughly $6 million to train compared to the estimated $100 million it cost to train GPT-4."
Julian Goldie @youtube_JulianGoldieSEO · 2026-04-09 view on x
"Deep Seek's engineers rewrote core parts of V4's code specifically for Huawei silicon. They gave Huawei early testing access and froze out US chip makers entirely."
Julian Goldie @youtube_JulianGoldieSEO · 2026-04-09 view on x
"SWE Bench Pro, 58.4. That is a new state-of-the-art, beating both GPT 5.4 and Claude Opus 4.6."
GPTAIclips Narrator @GPTAIclips · 2026-04-09 view on x
"SWE Bench verified at 78.8. Built on chain of thought that stays focused across hundreds of agent steps."
GPTAIclips Narrator @GPTAIclips · 2026-04-09 view on x
"The 31B dense model beats models 20 times its size on benchmarks. Supports text, images, audio, and video."
GPTAIclips Narrator @GPTAIclips · 2026-04-09 view on x
"Competitive benchmarks with full precision 8B models at 14 times less memory."
GPTAIclips Narrator @GPTAIclips · 2026-04-09 view on x
"On the OpenAI side, the company has been heavily teasing their new Spud model, actually doing more to hype it up than to tamp down expectations, reversing the trend that they've had all the way since back when GPT-5 underperformed."
AI Daily Brief Host @TheAIDailyBrief · 2026-04-09 view on x
"On Humanity's Last Exam, Opus got a 40% on a no-tools run compared to Mythos preview's 56.8%. With tools enabled, performance jumped to 64.7% compared to 53.1% for Opus."
AI Daily Brief Host @youtube_aidailybrief · 2026-04-09 view on x
"On OSWorld, which measures agentic computer use, Opus 4.6 got a 72.7% which jumped to 79.6% for Mythos."
AI Daily Brief Host @youtube_aidailybrief · 2026-04-09 view on x
"Anthropic ran the benchmark again using improvements from Terminal Bench 2.1 and extending the timeout window to 4 hours. Under those conditions, Mythos scored not 82% but 92.1%."
AI Daily Brief Host @youtube_aidailybrief · 2026-04-09 view on x
"The most forbidden technique is training an AI using interpretability techniques. You train on the final output only. Never the method. If you train on the method, you are training the AI to obfuscate its thinking."
Z @z_fortstrength · 2026-04-09 view on x
"Mythos appears to be the first class of models trained at scale on Blackwells. Reinforcement learning works and there is so much computing coming online soon."
Martin Casado @martin_casado · 2026-04-09 view on x
"I found more bugs in the last few weeks with Mythos than in the rest of my entire life combined."
Nicholas Caridi @nicholas_caridi · 2026-04-09 view on x
"Our ability to find weaknesses skyrocketed, but our ability to fix weaknesses didn't change. General generic toughening of code is a harder computer science problem than finding one vulnerability."
Wes Roth @WesRoth · 2026-04-09 view on x
"This is reflected in the high ability of LLMs to find security vulnerabilities versus the lower abilities of LLMs to rewrite whole systems with flawless security. I think people are thinking that Mythos autonomously finding bugs is the same as Mythos autonomously patching bugs and those are not the same thing."
Wes Roth @WesRoth · 2026-04-09 view on x
"Eight out of eight models detected Mythos flagship FreeBSD exploit. Because small, cheap, fast models are sufficient for much of the detection work, you don't need to judiciously deploy one expensive model and hope it looks in the right places."
Wes Roth @WesRoth · 2026-04-09 view on x
"Cloud Mythos is their most aligned model - good news, right? But also it can do the most damage if it's unaligned. It's like a less chance of doing something bad, but the capability is through the roof."
Wes Roth @WesRoth · 2026-04-09 view on x
"Mythos is very powerful and should feel terrifying. I am proud of our approach to responsibly preview it with cyber defenders rather than generally releasing it into the wild."
Boris Chernyy @boris_chernyy · 2026-04-09 view on x
"Mythos from Anthropic is another clear reminder that there is absolutely no wall in model capability progress right now. Meaningful double-digit gains on critical benchmarks."
Aaron Levie @aaronlevie · 2026-04-09 view on x
"In SWE-bench Pro for example, Claude Mythos beats out Opus 4.6 by 25%."
AI Explained @youtube · 2026-04-08 view on x
"Claude Mythos getting 83% beats out Gemini 3.1 Pro 82% and GPT 5.4 Pro at 80%. But on the subset remix, Claude Mythos gets the same score as Gemini 3.1 Pro and slightly underperforms GPT 5.4 Pro, which gets 88%."
AI Explained @youtube · 2026-04-08 view on x
"The geometric mean productivity uplift according to technical staff surveyed within Anthropic was 4x, four times the productivity when using Mythos."
AI Explained @youtube · 2026-04-08 view on x
"Humanity's Last Exam designed to test topics so obscure that it would indeed be the last exam that AI would saturate. Well, when allowed some tools, Claude Mythos gets almost two-thirds of those questions right compared to around 50% for other Frontier models."
AI Explained @youtube · 2026-04-08 view on x
"On multiple measures of software engineering, Mythos beats out Opus 4.6 by a massive margin. In SWE-bench Pro for example by 25%."
AI Explained @@aiexplained · 2026-04-08 view on x
"On the remix where you try to avoid memorization, Claude Mythos gets the same score as Gemini 3.1 Pro and slightly underperforms GPT 5.4 Pro, which gets 88%."
AI Explained @youtube_ai_explained · 2026-04-08 view on x
"Anthropics say it is not yet capable of causing dramatic acceleration. And yes, for followers of this channel, they admit that the previous survey they relied on for the release of Opus 4.6 was deeply flawed."
AI Explained @youtube_ai_explained · 2026-04-08 view on x
"Depending on whether you anchor on Claude Opus 4.5 or Claude Opus 4.6, one would nevertheless have to conclude that things are improving at an accelerating rate."
AI Explained @youtube_ai_explained · 2026-04-08 view on x
"It affected Claude Opus 4.6 and Sonnet 4.6 as well. When the reward code saw misaligned chains of thought, bad thoughts in other words, it could give a negative reward."
AI Explained @youtube_ai_explained · 2026-04-08 view on x
"With adaptive thinking, maximum effort, and Python tools, Claude Mythos preview scored almost 93% on finding specific UI elements in high-resolution screenshots. That is 10% higher than Claude Opus 4.6."
AI Explained @youtube/@aiexplainedYT · 2026-04-08 view on x
"75% - that is the score a new AI agent called GPT 5.4 just hit on the OS World benchmark. And look at that bar right next to it. That is the average human score sitting at 72.4%. The AI did not just match us, it actually beat us."
Dukta Feelgood @DuktaFeelgood · 2026-04-08 view on x
"It is a super robust simulation of the real everyday stuff that millions of us do at our computers every single day. We are not talking theory here. This is literally a simulation of professional white collar work. And the AI that pulled this off, it is not your standard chatbot. We are talking about GPT 5.4. It has got this massive 1 million token context window."
Dukta Feelgood @DuktaFeelgood · 2026-04-08 view on x
"The moment an AI can score higher than a person on real world tasks and do them on its own, well, we are not just getting close to that cliff anymore, we are standing right on the edge looking down."
Dukta Feelgood @DuktaFeelgood · 2026-04-08 view on x
"These models, GPT 5.4, Claude, Gemini, they are already available. So the advantage, the moat, it is not about having access to the tech. No, the real moat is implementation."
Dukta Feelgood @DuktaFeelgood · 2026-04-08 view on x
"Mythos scores 93% on this benchmark, and the current best public model, Opus, scores 80.8%. Nothing from OpenAI, Google, or any of the open source models are coming within 13 points of this."
josh @josh · 2026-04-08 view on x
"They pointed Claude Mythos at a Linux machine and without any advanced prompting, any advanced directives, it on its own found multiple kernel issues that it was able to chain together and just have root access to any system."
josh @josh · 2026-04-08 view on x
"They are officially declaring that there will be no public access to Claude Mythos until way more safeguards are in place."
josh @josh · 2026-04-08 view on x
"Mythos found a bug that had been sitting in the codebase for 27 years. It was a remote crash vulnerability where you can actually just connect to it and just shut it off."
josh @josh · 2026-04-08 view on x
"On SWE-bench Pro, Mythos got a 78% when previously Opus only got a 53. And if you are curious, I found the numbers for GPT 5.4 it was a 57.7. A 24 point jump is a 50% improvement on one of the hardest software benches we have."
Theo @t3dotgg · 2026-04-08 view on x
"They also massively increased their terminal bench score to an 82% previously at 65."
Theo @t3dotgg · 2026-04-08 view on x
"GPQA got from 91 to 94. They were always a little behind on this. So, yeah, I think it is pretty saturated at this point."
Theo @t3dotgg · 2026-04-08 view on x
"Humanity last exam, they went from a 40% to a 56.8%. And when given tools, it did even better at a 64.7%. Crazy that HLE is going to be saturated soon."
Theo @t3dotgg · 2026-04-08 view on x
"Mythos is to Opus what Opus is to Sonnet. It is a much bigger model that will be much more expensive that is slow but powerful and the capabilities are immense."
Theo @t3dotgg · 2026-04-08 view on x
"In our testing, Claude Mythos preview demonstrated a striking leap in cyber capabilities relative to prior models, including the ability to autonomously discover and exploit zero-day vulnerabilities in major operating systems and web browsers."
Theo @t3dotgg · 2026-04-08 view on x
"Mythos autonomously found and chained together several vulnerabilities in the Linux kernel and it allowed an attacker to escalate from an ordinary user to complete control of the machine. Finding a novel Linux exploit to get root is horrifying."
Theo @t3dotgg · 2026-04-08 view on x
"It does kind of suck that we are now at a place where there is a model that is 50% plus better than anything else out there that you can only use if you are on Anthropic nice guy list."
Theo @t3dotgg · 2026-04-08 view on x
"Anthropic has committed up to 100 million in usage credit for Mythos preview across these efforts as well as 4 million in direct donations to open source security organizations."
Theo @t3dotgg · 2026-04-08 view on x
"The security side seems to be an emergent behavior from getting good at code. They were not trying to train it to be good at hacking. They were just trying to make it good at code. And this just happened as a result."
Theo @t3dotgg · 2026-04-08 view on x
"They did not train the model to be good at cyber security. They trained it to be good at code. So if you can get good code and good code chat histories out of the other models and then RL an open-weight one, if you can get an open-weight model that is good enough at coding, it will also be able to pwn in a similar way."
Theo @t3dotgg · 2026-04-08 view on x
"The window between a vulnerability being discovered and being exploited by an adversary has collapsed. What once took months now happens in minutes with AI."
Theo @t3dotgg · 2026-04-08 view on x
"Claude Mythos preview is on essentially every dimension we can measure the best aligned model that we have released to date by a significant margin. Even so, we believe that it likely poses the greatest alignment related risk of any model we have ever released to date."
Theo @t3dotgg · 2026-04-08 view on x
"The SWE-bench multimodal implementation is nearly double as well. It is actually a bit over."
Theo @t3dotgg · 2026-04-08 view on x
"The price for Mythos preview is 25 dollars per million tokens in and 125 per million out. For reference, GPT-5.4 is 2.50 per million in and 15 per million out. So, it is approximately 10x more expensive than GPT-5.4."
Theo @t3dotgg · 2026-04-08 view on x
"The 31 billion parameter version of Gemma 4 is scoring in the same ballpark as models like Kimi K2.5 thinking. But here is the absurd part. I can run Gemma 4 locally with a 20 GB download, getting roughly 10 tokens per second on a single RTX 4090."
Jeff Delaney @@Fireship · 2026-04-08 view on x
"Some of the Gemma models have an E in the model name like E2B and E4B. That stands for effective parameters because these models incorporate per layer embeddings, which is like giving every layer in the neural network its own mini cheat sheet for each token."
Jeff Delaney @@Fireship · 2026-04-08 view on x
"The big model is small enough to run on a consumer GPU, and the Edge model is small enough to run on your phone or Raspberry Pi, while hitting intelligence levels that are on par with other open models that would normally require data center caliber GPUs just to run."
Jeff Delaney @@Fireship · 2026-04-08 view on x
"On SWE-bench Verified, Mythos scores at 93.9%, and Claude Opus 4.6, the model you can actually use right now, 80%. On Terminal Bench, Mythos actually hits 82% up from Opus's 65.4%."
TheAIGRID @theaigrid · 2026-04-08 view on x
"I was starting to believe that this kind of intelligence would not scale as the models got bigger because maybe you would have needed new architectures. But currently this new Mythos thing actually changes the game because it shows that there is a continued curve of improvement."
TheAIGRID @theaigrid · 2026-04-08 view on x
"Alongside Gemma 4, Google quietly dropped a research note on something called TurboQuant. It is a new approach to quantization that compresses model weights using polar coordinates and the Johnson-Lindenstrauss transform to shrink high-dimensional data down to single sign bits while preserving distances between data points."
Jeff Delaney @@Fireship · 2026-04-08 view on x
"Gemma 4 hits different because it is made in America, Apache 2.0 licensed, intelligent, and most importantly, tiny. Meta Llama models are quasi free and open under a special license that gives Meta leverage. OpenAI GPT OSS models are also Apache 2.0 but bigger and dumber than Gemma."
Jeff Delaney @@Fireship · 2026-04-08 view on x
"Software engineering bench verified, the real-world software engineering benchmark that everyone actually cares about. Mythos scores 93% on this benchmark, and the current best public model, Opus, scores 80.8%. Nothing from OpenAI, Google, or any of the open source models are coming within 13 points of this."
Josh @josh_youtube · 2026-04-08 view on x
"With adaptive thinking, maximum effort, and Python tools, Claude Mythos preview scored almost 93% on finding specific UI elements in high-resolution screenshots. That's 10% higher than Claude Opus 4.6."
AI Explained @aiexplained · 2026-04-08 view on x
"Nicholas Carini, a top cyber security expert, said using Mythos he's found more bugs in the last few weeks than in his entire career before that."
AI Explained @aiexplained · 2026-04-08 view on x
"To even double AI progress when you factor in compute being a key limiting ingredient, Anthropic predicts that you would actually need an uplift roughly 10 times larger than that 4x one. They think it would take a 40x productivity improvement to see a 2x progress speedup at Anthropic given the crazy bottleneck of compute."
AI Explained @aiexplained · 2026-04-08 view on x