Where was Mythos when WordPress fell? · patrik.re↓<br>Skip to main content
patrik.re
Table of Contents<br>Table of Contents
I am by no means an AI expert, and I get annoyed quickly by people who claim to be experts in the field, because let’s face it, at the end of the day we prompt that API and hope something useful comes out of it. When Anthropic announced the infamous Mythos model in spring of 2026 my entire feed on X was full of experts who were damn sure the Messiah had arrived and that cybersecurity was basically solved now. Anthropic released Mythos Preview to a handful of trusted entities under Project Glasswing, with the ultimate goal of “securing critical software for the AI era”. By June the whole thing had escalated to the point where the US government invoked export control powers and Anthropic shut off access to Fable 5 and Mythos 5 entirely.<br>In an update shared on 22nd of May, Anthropic stated: “For the last few months, Anthropic has used Mythos Preview to scan more than 1,000 open-source projects, which collectively underpin much of the internet—and much of our own infrastructure.”<br>Bold claims from the marketing department of Anthropic, leaving out one crucial aspect: which 1,000 projects were actually scanned by the almighty Mythos? Is it fair to assume that WordPress, that CMS from the 2000s that runs on roughly hundreds of millions of websites, wasn’t part of the open-source projects Anthropic looked at? Because surely it would’ve discovered the pre-authentication RCE in WordPress Core found by my colleague Adam Kues from Searchlight Cyber, right?<br>Anthropic will tell you that disclosed vulnerabilities are a lagging indicator, that they sit inside a 90 day window and can’t publish everything yet, and that’s a fair point to make. It just doesn’t cover this one. The WordPress 7.0.2 release notes credit the people who reported it, and Anthropic isn’t among them.<br>The numbers, if you read them<br>If you actually read the update, the numbers don’t quite add up. Mythos Preview found what Anthropic describes as 6,202 high or critical severity vulnerabilities out of 23,019 in total, and those severity ratings are the model’s own estimate of its own work. Of those 6,202, exactly 1,752 had been assessed by an independent security firm at the time of writing, and of the ones that got checked, a little over 60 percent held up as high or critical. So the number that kept popping up in my feed was a model grading its own homework, and the part where somebody else marked the homework covered about a quarter of it. Would you put that number in a press release?<br>Then there’s the question of what came out the other end. If you go looking for the public record of what Mythos actually found you land on roughly 40 CVEs, and 28 of those are Firefox and 9 of them are wolfSSL. A browser and a cryptography library. There’s no PHP in that list, no CMS, nothing that resembles the web most people actually use. Which brings me back to the same question: which 1,000 projects?<br>Some people actually checked<br>The people who did get their hands on it were nowhere near as excited as my timeline. Daniel Stenberg ran it against curl and got five reported findings, of which exactly one turned out to be a real vulnerability, and it was low severity, and his conclusion was that the hype around the model was “primarily marketing”. Bruce Schneier argues that it had become common wisdom that Mythos is better at finding software vulnerabilities than other models, “which is just not true.” Two people with nothing to sell, saying so in public, months ago. Did it change anything at all? Not that I noticed.<br>Securing critical software, in practice<br>And then there’s the part that bothers me the most. Anthropic reported around 530 high or critical bugs to open source maintainers, of which 75 had been patched by the time they published that update, and several maintainers told them they were so overloaded that they asked Anthropic to slow the disclosures down. So securing critical software for the AI era arrived, in practice, as a queue of tickets landing on unpaid volunteers who were already drowning.<br>Cloudflare published an account of what running this thing actually looks like. They tested it against more than fifty of their own repositories and found that a single agent session covers, in their words, “maybe a tenth of a percent of the surface in a useful way” before the context window fills up, so they ended up building an eight stage pipeline around the model to get anything worth reading out of it.<br>What Adam did, for the record, was point GPT5.6 Sol Ultra at a clean WordPress tree with the git history stripped out so it couldn’t cheat by reading patches, cap it at four agents, and leave it running. Ten hours later he had the chain, and it cost him about 25 dollars. His own conclusion is that no security...