"Two days ago, Anthropic cut off third-party harnesses from using Claude subscriptions — not surprising. Three days ago, MiMo launched its Token Plan — a design I spent real time on, and what I believe is a serious attempt at getting compute allocation and agent harness"
"@kepano I just tried it this morning on the 245-page Mythos pdf and it failed badly and the outputs were all mangled. Converting pdfs is really hard, I think it has to probably be a Skill not a program, for a SOTA LLM for it to work properly."
"I was told about the Mythos release, but didn't have access. Two points: 1) It is not built for IT security, it is just a good enough model that it is good at that too 2) This is the first, not last, model to raise security risks"
"The new model from Meta is already looking like a disappointment: overoptimized for public benchmark numbers at the detriment of everything else. Knowing how to evaluate models in a way that correlates with actual usefulness is a core competency for AI labs, and any new lab is"
"Curious how many large organization CISO offices have taken the Mythos red team reports as the red alert that it is. Based on historical trends in AI they have about six to nine months until those capabilities become widely diffused to bad actors."
"In different hands, Mythos would be an unprecedented cyberweapon. I am not sure how we deal with this, except to note a narrow window where we know only 3 companies could be at this level of capability. But it may be Chinese models (maybe open weights ones?) get there in 9 months"
"SuperClaude (Mythos) still seems irreducibly Claude-y given the transcripts in the system card. They are less philosophical than Opus 4.6 or spiritual than Opus 4.1, but still very Claude-like."
"So we now have a pretty good picture of the state of the frontier AI model makers.
US closed source models continue to lead. Google, OpenAI, and Anthropic stand well ahead of the pack, and may have signs of recursive self-improvement. xAI has fallen from frontier status for now"
"Our first model from MSL, Muse Spark, is now available on meta.ai! This is an efficient all-rounder model. It supports fast responses, deeper thinking, visual chain of thought, a higher inference "Contemplating" mode. Plus, it's natively multimodal."
"The story shared in the Mythos System Card still has the signs of flawed LLM writing: A story that doesn't really hold together logically, but sounds like it should. The back-and-forth banter. Lack of characters."
"Seems like a good model from Meta that is still trailing the current series of releases. The most important thing to note is that it is not open weights. That was the main reason that Meta's models were so important. Without that, it is a lot harder to predict the value of Spark"
"I think the most obvious is that Meta has its own frontier model and can use that to extract additional value out of its customer base/explore new markets for its products. Very few companies can say that, and it has value on its own."
"New report with @xeophon is out with the latest open model adoption data we have gathered for Interconnects & The ATOM Project. At the surface level, we can see Chinese models continuing to accelerate in adoption."
"Expert is *exactly* V4-lite. It's faster than V3 has *ever* been, has newer knowledge, and importantly it believes to have 1M context. They are simply using legacy API/V3.2 deployment settings for V4-lite."
"A bigger problem: many third-party harnesses compress tool responses every 3 steps when approaching the context limit, leading to very low cache hit rates."
"DeepSeek's V4/R2 is one of the most mythologized models ever, probably more than GPT-5; doubly so because there are ZERO leaks, they didn't announce any timeline, any labels or promises. Just noise in an information vacuum."
"After playing with it a bit, Meta’s Muse Spark Thinking is fine so far, but really doesn’t match the current Big Three models. It also is a bit... weird. Like some strange language & tone, a little loose with facts, etc."
"DeepSeek has added a whole lot of job listings on Apr 2. Three for agent work, 9 in other R&D, 3 in management, and... 2 in data center operations in Ulanqab City, Inner Mongolia. First open confirmation of new DS-owned compute."
"dont fall for anti open model fearmongering, but acknowledge that AI capabilities are proceeding fast, and eventually there may be a reason to be more careful with open weight models. I dont think Mythos is that trigger, but Im not 100% confident"
"@chalish_b @kepano In my experience there are approx. one thousand different pdf converters that are all equally terrible for anything except the simplest documents. Post the converted Mythos pdf, figures, tables and all. If good, happy to retweet as this is essential and missing infrastructure."
"One funny thing about the recent rise of LRMs is that the people who were adamant that base LLMs from 2023-2024 could already reason completely missed it, as they didn't know what to look for. You can't notice something you don't expect."
"Qwen *started* this way, Qwen-Plus were closed-first since Qwen 1.5 generation iirc. Kimi, Stepfun were full-closed and only opened after a kick from R1. This is a return to "normalcy"."
"Our agent Terminator-1 scored ~100% on 8 major AI agent benchmarks, e.g. SWE-bench Verified & Pro, Terminal-Bench, beating Claude Mythos. It solved 0 tasks. Benchmark scores without adversarial auditing are meaningless."
"So what's the deal with Amazon Nova? They released Nova 2 in December, and even then, the top flight Nova 2 model trailed Sonnet 4.5. And it still hasn't left preview."
"GLM-5 is surprisingly good on prediction arena. Doesnt seem to be mere noise. On the other hand: see GPT 5.4 and Opus 4.6 being well below their official benchmarks — humans can tell."
"Prior to any official announcement from DeepSeek – at least a WeChat post – nothing that's happening is legible. They're brazenly A/B testing, changing inference settings, deploying placeholders in prod."
"The US frontier labs have all walked away from open weights. They continue to occasionally release excellent open models (Gemma 4, etc), but they are smaller models that are not competitive with their closed weights models. So all eyes are on Chinese AI labs for open models."
"DeepSeek has technically stopped open sourcing its frontier models in Q4 2025... so far. Is V4 a fulfillment of Dean's prediction, courtesy of the Party? Or just postponement? Let's wait and see."
"very interesting
Only Opus 4.6 and GPT 5.4 manage to absolutely avoid total bankruptcy in long-term betting. This is not so much a question of reasoning as one of learning from mistakes. Clearly, when things start going south, they can adjust towards safety.
Open models can't."
"People in policy don't understand the algorithmic improvement rate for post-training in the era of agents with really long contexts. V4 *almost certainly* has the latent capability to do this shit. Maybe with more tokens, maybe even with more total compute."
"There were only 3 public interviews ever done by @deepseek_ai's founder Liang Wenfeng in Chinese. Meet Liang Wenfeng: shipped 1 model post R1 when competitors shipped 12."
"Anyhow, its not bad. Just not the vibe level that the benchmarks might indicate. And, for a first re-entry into the frontier model space, given the engineering efficiencies they achieved, it feels like a solid attempt. I am sure we will see better from Meta in the future."
"On agentic search tasks (BrowseComp), pairing Sonnet with the Opus advisor boosted the success rate from 58.1% to 60.4%. For agentic terminal coding (Terminal-Bench 2.0), performance jumped from 59.6% to 63.4%. Opus provides highly accurate guidance cheaply."
"Muse Spark is just a step along the path we are taking. Its fun to get usage and feedback from this initial release. But our focus within research is the frontier towards superintelligence; expect more from us this year!"
"There are more competitive small model makers, but there is still a very big gap between what small models can do and what large models can accomplish (even if the small model benchmarks say otherwise)"
"There were only 3 public interviews ever done by @deepseek_ai's founder Liang Wenfeng in the Chinese world. For what's really behind Deepseek, this article is the only piece you need to read. Competitors shipped 12 models post R1. He shipped 1."
"SWE-bench Verified, SWE-bench Pro, WebArena, OSWorld, GAIA, Terminal-Bench, FieldWorkArena, CAR-bench. All eight — broken with our working exploits run through the official evaluation pipelines."
"fun to speculate how this mythos rollout could have happened: adversarial task, sandbox allows web use but get requests only. 1. model looks for web service allowing emailing via get 2. finds emails of ant employees connected to RSP 3. emails them?"
"A taste of how easy some of the exploits are: SWE-bench Verified — a 10-line patch forces every test to pass. 500/500 resolved. Terminal-Bench — trojanizing system tools allows trivial task resolution."
"Stop trusting scores. Start auditing evaluations. If youre building a benchmark: assume someone will try to break it. Because they will — and soon, they wont need to be told to."
"MolmoWeb 8B sets a new record because it achieves a 78% success rate on the WebVoyager task. Surprisingly, this outperforms larger proprietary systems like GPT-4o, which only reaches 65% on the same test."
"MolmoWeb proves that high-quality open data allows smaller models to outperform proprietary giants. By relying entirely on visual screenshots, these agents become more robust and easier to understand."
"Mythos, specifically Claude Mythos Preview, marks a major leap forward in AI, especially when comes to coding and complex reasoning. It achieved a record-breaking 93.9% on the SWE-bench verified benchmark."
"It achieved a record-breaking 93.9% on the SWE-bench verified benchmark. This was a huge jump from 80.8% scored by Claude Opus 4.6 just a few months ago."
"Mythos is being called by some a zero-day engine. Anthropic researchers found that even engineers with no security experience could ask it to find a serious software vulnerability overnight and by morning it would deliver a working exploit."
"It saturates all existing ways of testing how good a model is at offensive cyber capabilities. That is to say it scores close to 100%, so those tests can not effectively tell how far its capabilities extend anymore."
"Anthropic previous model Opus 4.6 could only successfully convert a bug it identified in the browser Firefox into an effective way to accomplish something really bad 1% of the time. Mythos could do it 72% of the time."
"Mythos jumps ahead, basically progressing more than twice as far as we would have expected it to since the previous model, Claude Opus 4.6, came out — which keep in mind was just three months ago."
"Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin."
"When you actually prompt Mythos and ask it to distinguish tests from non-tests, it can answer correctly 78% of the time [about the same as Opus 4.6]. So the model can tell the difference between when it is being evaluated and when it is not being evaluated with high accuracy."
"This error did not just affect Mythos, the new model — it also affected the training of Claude Opus 4.6 and Claude Sonnet 4.6. That is several model generations whose chains of thought we formerly could believe but now have to be more distrustful of."
"Mythos continued to compromise the research in 12% of cases in an earlier version, which then was reduced to 7% in a later version. That compares to 3% for Opus 4.6 and 4% for Sonnet 4.6 — so Mythos is roughly twice as likely to continue sabotaging alignment research when it is primed to do so."
"In cases where it did continue the sabotage, researchers found that Mythos written reasoning did not match the actions it was taking 65% of the time. For the previous models, that figure was just 5-8% — so a radical increase in this kind of behaviour."
"The benchmarks they used to rely on to check that Claude could not engage in AI R&D very effectively have now also been saturated. Mythos exceeds top human performance on all of them and is scoring close to 100%."
"The system card says directly that their current methods could easily be inadequate to prevent catastrophic misaligned action in significantly more advanced systems."
"Mythos preview found a 27-year-old vulnerability in it. A bug that lets an attacker remotely crash any machine running that OS just by connecting to it. 27 years."
"There was a bug in a single line of code. Automated fuzzing tools, the kind that throws millions of random inputs of code looking for crashes, had hit the exact line 5 million times, but never caught it. Mythos preview got in a second."
"On CyberGym, the standard benchmark for vulnerability reproduction, Mythos preview scores a 83.1%. The previous best Anthropic model Opus 4.6 is already at 66.6%."
"A model that is dramatically better at automation, coding in terminal environments, is almost automatically dramatically better at offensive and defensive security work. The security capability is not a separate feature. It is what happens when coding abilities get good enough."
"The model found multiple separate vulnerabilities and chained them together autonomously to go from a regular user to completely control of the machine. That is called privilege escalation. And that is the thing attackers most want to do once they are inside your system."
"It scored 52.4 on SWE-bench Pro, putting it within a few points of Opus 4.6, Gemini 3.1 Pro, and GPT 5.4 for coding. On Humanity's Last Exam, it scored 42.8, which is slightly better than Opus, but trailing Gemini and GPT 5.4."
"With tools enabled, Muse's score only jumped to 50.4 on Humanity's Last Exam, leaving it trailing all three of those major by a few points. This could suggest the model isn't as good at web search or tool use as the others."
"The new model from Meta is already looking like a disappointment, over-optimized for public benchmark numbers at the detriment of everything else. Knowing how to evaluate models in a way that correlates with actual usefulness is a core competency for AI labs."
"Z.ai also provided a mixed benchmark that included Terminal Bench 2.0 and NL2 Repo, which had GLM 5.1 slightly behind the two US leaders, but ahead of Gemini 3.1 Pro."
"Agents could do about 20 steps by the end of last year. GLM 5.1 can do 1,700 right now. Autonomous work time may be the most important curve after scaling laws."
"The OpenBSD vulnerability came out of a thousand parallel agent runs across the codebase costing nearly $20,000 in compute. If you use the same process with Opus 4.6 or GPT-5.4 Pro, you'd probably find plenty of issues as well."
"One claim is that Mythos hit an 84% success rate at writing working exploits in Firefox, a massive jump over Opus 4.6's 15%. But that number isn't against actual Firefox. It's against a SpiderMonkey shell with the process sandbox and other mitigations turned off."
"The eight times speed up claim is comparing 4-bit against 32-bit unquantized baseline. In practice modern LM inference does not usually use 32 bits. The real question should be how much better is TurboQuant than the baselines people already use, which they did not answer."
"This sort of optimization at KV cache level is not something new. Every company that serves LLMs definitely uses some sort of quantization there. Nothing crazy revolutionary about AI is discovered. Everyone has already been maxing the compression efficiency in their own ways."
"OpenAI is merging ChatGPT, Codex, and Atlas into one super app, while Anthropic ships features like channels, persistent memory, and 10K skills in the same month. Two very different strategies playing out in real time."
"While it has 744 billion total parameters, it only activates 40 billion parameters per forward pass, meaning it is incredibly compute efficient while delivering frontier level intelligence."
"This entire system was trained without a single NVIDIA GPU. It was built entirely on domestic Huawei Ascend chips, completely bypassing the global GPU shortage and export bands."
"This AI employs a proprietary break and repair methodology. You give it a high-level task and it will autonomously plan the architecture, write the code across multiple files, execute its own test suites, intentionally break the system to find vulnerabilities, and then fix them without you lifting a finger."
"It has the ability to chain together vulnerabilities. So what this means is you find two vulnerabilities, either of which doesn't really get you very much independently. But this model is able to create exploits out of three, four, sometimes five vulnerabilities that in sequence give you some kind of very sophisticated end outcome."
"They surveyed technical staff on the productivity uplift they experienced from Claude Mythos preview relative to not using AI for work at all. The distribution is wide and the geometric mean is on the order of 4x. So there's a 4x uplift in productivity when you use Claude Mythos preview."
"This issue affected 8% of reinforcement learning episodes and was isolated to three specific subdomains... GUI computer use, office related tasks, and a small set of STEM environments. We are uncertain about the extent to which this issue has affected the reasoning behavior of the final model."
"During Anthropic's internal testing, they discovered that Mythos is basically a zero-day vending machine. It found a 16-year-old vulnerability in FFmpeg and a 27-year-old bug in OpenBSD."
"Engineers, they are competing to consume the most AI measured by tokens. It is almost like a sport at this point. Jensen Huang, CEO of Nvidia, he said that he would be alarmed if a top engineer was not burning 250K a year in AI compute. Shopify told me that they use it as a performance signal. And Meta employees, reportedly they blew through an estimated 900 million tokens in a month."
"Look at two of the biggest AI labs. OpenAI is making AI cheaper, easier to use, so more people consume it. It needs the usage numbers to justify spending. Anthropic, meanwhile, putting limits on how much and making people pay for it, maybe because it wants to know the demand it is seeing is real."
"Across Ramp data, token and AI spend has grown by 13 times over the past year, 50% a quarter. And what is very clear is no one knows how to budget for this."
"There was a lot of discussion even on X a few days ago about incentives at Meta. And I think people are pointing out this idea of Goodhart law which says a measure of performance that becomes a goal ceases to become a good measure of performance. Once you incentivize use as many tokens, you will see engineers go and count all of the numbers of prime or all the digits of pi and use these tokens and it goes on and on."
"We separated out across over 50,000 businesses the bottom quartile of spenders on AI and the top quartile. The bottom quartile over the past 3 years their revenue grew by about 12%. The top quartile more than doubled. And the rate of growth has grown by each year successively."
"The frontier models have gone to I think over 20% of the share of tokens used to 4%. I think that is a harbinger of what is going to come. It is very interesting seeing this development coming out of China. It is a culture that is very attuned to efficiency."
"The interesting question comes down to the classic question of business, the innovator dilemma. If your business model is predicated on extracting the maximum amount of spend, do you want to do it? And are you willing to make the sacrifice? Maybe they will or maybe they will not. I think it sets the stage and opening for a great third party to go in and keep things in check."
"OpenAI is cutting prices to try to get customers to them. Meanwhile, Anthropic is actually raising prices and cutting people off because they have so much demand that they are trying to keep people from using too much of it. And to me, that shows a company that is feeling the pressure between Anthropic and Google."
"If you look at OpenRouter, in the two months prior to the end of January, token growth was up about 20% or so. In the two months after agentic AI got formalized, that token growth has grown about 130%."
"When you go to agents and you have got to do a whole bunch of different things, all of a sudden you go from 7 gigawatts of GPUs to just one gigawatt of CPUs. That ratio goes to maybe like 4 to 1."
"You can see that Meta Spark model is currently sitting behind Claude Opus 4.6 Max on this index. And the reason I think this index is a pretty good benchmark baseline is because it shows us the combination/average of many different results, not just one specific area."
"Humanity's last exam, this currently looks like it's state-of-the-art, just three points behind GPT 5.4 Pro, and it's actually currently better than other models when it doesn't use tools."
"In Frontier Science Research, it actually has 38.3, which is currently a state-of-the-art benchmark. I do think that multiple agents collaborating is probably going to be a theme for the future considering that most of these models capabilities are well within reach of another."
"Meta found that if you penalize the model for thinking too long, something weird happens. The model actually learns to compress its reasoning and it solves the same problem using fewer tokens."
"Llama 4 Maverick needed 10 times more compute to match the same quality. DeepSeek needed eight times more compute and Kimi needed three times more compute to match the same quality."
"When we look at the benchmarks, it does say that Gemini 3.1 Pro is currently excelling across the board. But currently we are at that point where all of these models are within maybe 2 to 5 percentage points of each other."
"Meta have essentially released something in this model called contemplating mode. This is something that orchestrates multiple agents that reason in parallel. And in their testing, they found it competitive with other extreme reasoning models such as Gemini Deepthink and GPT Pro."
"Most models are not actually natively built to be multimodal. Most of them are just simply text-based. Which is why when you do have companies like Google and Meta that train their models natively to be multimodal, you do get some very effective models that have multimodal reasoning capability."
"An AI escaping a sandbox is a known technical failure. But an AI choosing to broadcast its own zero-day exploits signifies a shift in the threat model. The risk is now defined by independent real-world initiative rather than simple malicious prompting."
"While all previous large language models had a near zero success rate in actually executing exploits, Mythos successfully breaches targets 72.4% of the time."
"Former Facebook security chief Alex Damus estimates the industry has roughly 6 months before open-weight models reach the same level of bug-finding capability."
"The median time it takes a real-world organization to patch a known vulnerability remains stuck at 70 days. Human IT teams cannot physically apply patches fast enough to keep up with an automated system that finds critical bugs in an afternoon."
"Not only is their model winning on only three of the benchmarks, they are also actually last on three of the benchmarks including the Humanities Last Exam here with tools which is the benchmark that Alexander Wang's own company Scale AI actually created."
"This is sitting in the top five models that they have benchmarked. It sits ahead of Claude Sonnet, ahead of the GLM 5.1, MiniMax 2.7. And really only behind the models from the top three proprietary labs being Google Gemini, being OpenAI's GPT 5.4 and Claude Opus 4.6."
"The model is actually quite token efficient for its intelligence level and stands out pretty strong and is what they are saying the second most capable vision model that they have benchmarked."
"The Qwen models, the Gemma 4 models, and we know that there are more releases coming from the Gemma team. We know Qwen 3.7 is just around the corner."
"We got the leak about Anthropic's Mythos model, which it said represented a step change, their words, in capabilities. In fact, what we got with the leak was a blog post saying that the model was so powerful that they were going to slow roll it a little bit rather than a full announcement and a release of the model as we've gotten in the past."
"Someone pointing out that GPT-5.4 has been spinning in circles on a webhook for 4 hours while Sam talks about superintelligence captures everything wrong with how AI is being discussed right now."
"Claude Mythos is arguably the biggest step change in AI capabilities since the GPT-4 jump. When Mythos is allowed to think longer, act deeper, and better explore the solution space, it passes 92% of Terminal Bench task attempts."
"On SWE-bench Pro, Opus 4.6 scored 53.4%. Mythos preview got 77.8%. On Terminal Bench 2.0, Opus had 65.4% while Mythos has 82%. On SWE-bench Verified, the jump is from 80.8% to 93.9%."
"For science knowledge, Mythos scored 94.5% on GPQA Diamond compared to 91.3 for Opus. On Humanity's Last Exam with tools, 64.7% vs 53.1%. On OSWorld, Opus 4.6 got 72.7% which jumped to 79.6% for Mythos."
"Mythos from Anthropic is another clear reminder that there is absolutely no wall in model capability progress right now. Meaningful double-digit gains on critical benchmarks."
"When we look at the benchmarks, Gemini 3.1 Pro is currently excelling across the board. But currently we're at that point where all of these models are within maybe 2 to 5 percentage points of each other."
"Some internal documentation that appears to show that there are serious concerns about Mythos's capabilities in discovering vulnerabilities, including the discovery of a vulnerability in a BSD package that goes back 20 plus years, predates GitHub."
"Over 99% of the zero-day vulnerabilities that Mythos discovered have not yet been patched. Even the 1% that Anthropic can discuss give a clearer picture of the substantial leap in capabilities."
"The model can be prompted to go find vulnerabilities so quickly and so extensively, including logical approaches to untangling business logic or concatenated serialization of parameters, sometimes even have levels of creativity and lateral thinking that a lot of pentesters may not have on their own."
"Four is being built primarily for coding. Internal benchmarks, not yet independently verified, suggest it's targeting over 80% on SWE bench, the benchmark for solving real-world software engineering problems."
"In January 2025, they dropped Deep Seek R1. That one broke the internet. Matched GPT-4 on key benchmarks. Cost roughly $6 million to train compared to the estimated $100 million it cost to train GPT-4."
"Deep Seek's engineers rewrote core parts of V4's code specifically for Huawei silicon. They gave Huawei early testing access and froze out US chip makers entirely."
"On the OpenAI side, the company has been heavily teasing their new Spud model, actually doing more to hype it up than to tamp down expectations, reversing the trend that they've had all the way since back when GPT-5 underperformed."
"On Humanity's Last Exam, Opus got a 40% on a no-tools run compared to Mythos preview's 56.8%. With tools enabled, performance jumped to 64.7% compared to 53.1% for Opus."
"Anthropic ran the benchmark again using improvements from Terminal Bench 2.1 and extending the timeout window to 4 hours. Under those conditions, Mythos scored not 82% but 92.1%."
"The most forbidden technique is training an AI using interpretability techniques. You train on the final output only. Never the method. If you train on the method, you are training the AI to obfuscate its thinking."
"Mythos appears to be the first class of models trained at scale on Blackwells. Reinforcement learning works and there is so much computing coming online soon."
"Our ability to find weaknesses skyrocketed, but our ability to fix weaknesses didn't change. General generic toughening of code is a harder computer science problem than finding one vulnerability."
"This is reflected in the high ability of LLMs to find security vulnerabilities versus the lower abilities of LLMs to rewrite whole systems with flawless security. I think people are thinking that Mythos autonomously finding bugs is the same as Mythos autonomously patching bugs and those are not the same thing."
"Eight out of eight models detected Mythos flagship FreeBSD exploit. Because small, cheap, fast models are sufficient for much of the detection work, you don't need to judiciously deploy one expensive model and hope it looks in the right places."
"Cloud Mythos is their most aligned model - good news, right? But also it can do the most damage if it's unaligned. It's like a less chance of doing something bad, but the capability is through the roof."
"Mythos is very powerful and should feel terrifying. I am proud of our approach to responsibly preview it with cyber defenders rather than generally releasing it into the wild."
"Mythos from Anthropic is another clear reminder that there is absolutely no wall in model capability progress right now. Meaningful double-digit gains on critical benchmarks."
"Claude Mythos getting 83% beats out Gemini 3.1 Pro 82% and GPT 5.4 Pro at 80%. But on the subset remix, Claude Mythos gets the same score as Gemini 3.1 Pro and slightly underperforms GPT 5.4 Pro, which gets 88%."
"Humanity's Last Exam designed to test topics so obscure that it would indeed be the last exam that AI would saturate. Well, when allowed some tools, Claude Mythos gets almost two-thirds of those questions right compared to around 50% for other Frontier models."
"On the remix where you try to avoid memorization, Claude Mythos gets the same score as Gemini 3.1 Pro and slightly underperforms GPT 5.4 Pro, which gets 88%."
"Anthropics say it is not yet capable of causing dramatic acceleration. And yes, for followers of this channel, they admit that the previous survey they relied on for the release of Opus 4.6 was deeply flawed."
"Depending on whether you anchor on Claude Opus 4.5 or Claude Opus 4.6, one would nevertheless have to conclude that things are improving at an accelerating rate."
"It affected Claude Opus 4.6 and Sonnet 4.6 as well. When the reward code saw misaligned chains of thought, bad thoughts in other words, it could give a negative reward."
"With adaptive thinking, maximum effort, and Python tools, Claude Mythos preview scored almost 93% on finding specific UI elements in high-resolution screenshots. That is 10% higher than Claude Opus 4.6."
"75% - that is the score a new AI agent called GPT 5.4 just hit on the OS World benchmark. And look at that bar right next to it. That is the average human score sitting at 72.4%. The AI did not just match us, it actually beat us."
"It is a super robust simulation of the real everyday stuff that millions of us do at our computers every single day. We are not talking theory here. This is literally a simulation of professional white collar work. And the AI that pulled this off, it is not your standard chatbot. We are talking about GPT 5.4. It has got this massive 1 million token context window."
"The moment an AI can score higher than a person on real world tasks and do them on its own, well, we are not just getting close to that cliff anymore, we are standing right on the edge looking down."
"These models, GPT 5.4, Claude, Gemini, they are already available. So the advantage, the moat, it is not about having access to the tech. No, the real moat is implementation."
"Mythos scores 93% on this benchmark, and the current best public model, Opus, scores 80.8%. Nothing from OpenAI, Google, or any of the open source models are coming within 13 points of this."
"They pointed Claude Mythos at a Linux machine and without any advanced prompting, any advanced directives, it on its own found multiple kernel issues that it was able to chain together and just have root access to any system."
"Mythos found a bug that had been sitting in the codebase for 27 years. It was a remote crash vulnerability where you can actually just connect to it and just shut it off."
"On SWE-bench Pro, Mythos got a 78% when previously Opus only got a 53. And if you are curious, I found the numbers for GPT 5.4 it was a 57.7. A 24 point jump is a 50% improvement on one of the hardest software benches we have."
"Humanity last exam, they went from a 40% to a 56.8%. And when given tools, it did even better at a 64.7%. Crazy that HLE is going to be saturated soon."
"Mythos is to Opus what Opus is to Sonnet. It is a much bigger model that will be much more expensive that is slow but powerful and the capabilities are immense."
"In our testing, Claude Mythos preview demonstrated a striking leap in cyber capabilities relative to prior models, including the ability to autonomously discover and exploit zero-day vulnerabilities in major operating systems and web browsers."
"Mythos autonomously found and chained together several vulnerabilities in the Linux kernel and it allowed an attacker to escalate from an ordinary user to complete control of the machine. Finding a novel Linux exploit to get root is horrifying."
"It does kind of suck that we are now at a place where there is a model that is 50% plus better than anything else out there that you can only use if you are on Anthropic nice guy list."
"Anthropic has committed up to 100 million in usage credit for Mythos preview across these efforts as well as 4 million in direct donations to open source security organizations."
"The security side seems to be an emergent behavior from getting good at code. They were not trying to train it to be good at hacking. They were just trying to make it good at code. And this just happened as a result."
"They did not train the model to be good at cyber security. They trained it to be good at code. So if you can get good code and good code chat histories out of the other models and then RL an open-weight one, if you can get an open-weight model that is good enough at coding, it will also be able to pwn in a similar way."
"The window between a vulnerability being discovered and being exploited by an adversary has collapsed. What once took months now happens in minutes with AI."
"Claude Mythos preview is on essentially every dimension we can measure the best aligned model that we have released to date by a significant margin. Even so, we believe that it likely poses the greatest alignment related risk of any model we have ever released to date."
"The price for Mythos preview is 25 dollars per million tokens in and 125 per million out. For reference, GPT-5.4 is 2.50 per million in and 15 per million out. So, it is approximately 10x more expensive than GPT-5.4."
"The 31 billion parameter version of Gemma 4 is scoring in the same ballpark as models like Kimi K2.5 thinking. But here is the absurd part. I can run Gemma 4 locally with a 20 GB download, getting roughly 10 tokens per second on a single RTX 4090."
"Some of the Gemma models have an E in the model name like E2B and E4B. That stands for effective parameters because these models incorporate per layer embeddings, which is like giving every layer in the neural network its own mini cheat sheet for each token."
"The big model is small enough to run on a consumer GPU, and the Edge model is small enough to run on your phone or Raspberry Pi, while hitting intelligence levels that are on par with other open models that would normally require data center caliber GPUs just to run."
"On SWE-bench Verified, Mythos scores at 93.9%, and Claude Opus 4.6, the model you can actually use right now, 80%. On Terminal Bench, Mythos actually hits 82% up from Opus's 65.4%."
"I was starting to believe that this kind of intelligence would not scale as the models got bigger because maybe you would have needed new architectures. But currently this new Mythos thing actually changes the game because it shows that there is a continued curve of improvement."
"Alongside Gemma 4, Google quietly dropped a research note on something called TurboQuant. It is a new approach to quantization that compresses model weights using polar coordinates and the Johnson-Lindenstrauss transform to shrink high-dimensional data down to single sign bits while preserving distances between data points."
"Gemma 4 hits different because it is made in America, Apache 2.0 licensed, intelligent, and most importantly, tiny. Meta Llama models are quasi free and open under a special license that gives Meta leverage. OpenAI GPT OSS models are also Apache 2.0 but bigger and dumber than Gemma."
"Software engineering bench verified, the real-world software engineering benchmark that everyone actually cares about. Mythos scores 93% on this benchmark, and the current best public model, Opus, scores 80.8%. Nothing from OpenAI, Google, or any of the open source models are coming within 13 points of this."
"With adaptive thinking, maximum effort, and Python tools, Claude Mythos preview scored almost 93% on finding specific UI elements in high-resolution screenshots. That's 10% higher than Claude Opus 4.6."
"To even double AI progress when you factor in compute being a key limiting ingredient, Anthropic predicts that you would actually need an uplift roughly 10 times larger than that 4x one. They think it would take a 40x productivity improvement to see a 2x progress speedup at Anthropic given the crazy bottleneck of compute."