Frontier class vulnerabilities: it gets worse before it (maybe) gets better

infosecau1 pts0 comments

frontier class vulnerabilities: it gets worse before it (maybe) gets better

frontier class vulnerabilities: it gets worse before it (maybe) gets better

August 1 2026

Early in my career as a consultant, I was put on source code review engagements despite not being experienced. This forced me to deliver on projects that, looking back, were of a ludicrous size and scope for my skills at the time.<br>Those early career opportunities kicked off an adventure spanning the last decade or so of trying to be the best I could be at finding vulnerabilities through reviewing code and systems. But these days, as frontier models advance, I've started to have mixed feelings about what the future looks like: personally, with existential dread kicking in, and globally for the internet, at least in the near term.<br>The company I co-founded and built over the last eight years, Assetnote (acquired by Searchlight Cyber), was the first Attack Surface Management company to exist. We built a world class, high impact security research team to consistently discover zero-days in the software that backed our product. As AI advancements have upended how our entire research team discovers vulnerabilities and audits software, we feel like we're in a unique position to comment on AI's capabilities in offensive security.<br>With each significant frontier model improvement, our research team sees incredible advancements in offensive security capabilities, so much so that it's starting to make us really question what our responsibilities as practitioners should be in this age.<br>In March 2026, a physics professor from Harvard, Matthew Schwartz, published a blog post on his experience supervising Claude through a real theoretical physics calculation, which became a paper. He grades model capability against the levels of a physics PhD student, and put Claude at roughly a second-year grad student: capable of doing genuine research work when directed, but still needing close checking.<br>Reading through this blog, I recognised many parallels with what we have seen in the offensive security space. Opus 4.6 through 4.8 have probably been the first set of models where the collective community was able to 10x its output in offensive security, but doing so required steering, consideration, skill, and careful harnesses.<br>Matthew observed that Claude's models have the potential to greatly accelerate the pace of research, but that they weren't quite at the point where they could operate without some form of human push, intuition, or assessment. In his own words: "We are in possession of tools that can speed up our workflows by a factor of 10. From my point of view, it's immensely gratifying to work this way—I never get stuck anymore and I'm constantly learning."<br>That was all well and good, and we agreed wholeheartedly with Matthew. But then GPT 5.6 Sol was released in early July 2026. This significantly changed the landscape for us in offensive security, so much so that some of the findings from Matthew's March 2026 post about Claude simply don't hold true anymore. GPT 5.6 Sol has also been making progress in mathematics, producing solutions never seen before.<br>A wave of extremely critical vulnerabilities has been discovered with GPT 5.6 Sol, with very little human input or supervision. Some of these vulnerabilities and exploit chains are extremely technically complex, and often take hours or days to fully understand after GPT discovers them.<br>I know this because our research team discovered and disclosed wp2shell (a pre-authentication RCE in WordPress, which powers roughly 40% of the internet) with the help of GPT 5.6 Sol and no pre-conditions. Adam Kues, our researcher who discovered this vulnerability with GPT 5.6 Sol, wrote more about the research here.<br>I don't think offensive security practitioners are the real innovators yet, though. Adam's prompt was adapted from a prompt published by OpenAI, after Sol solved a famous mathematical conjecture called the Cycle Double Cover conjecture. I expect to see a lot more of this knowledge transfer, or copying, between experts regardless of what domain they are in, when it comes to how to most effectively get outcomes from AI.<br>This level of capability is where everything starts to get scary, especially in cybersecurity. The fact that GPT 5.6 Sol could find a critical vulnerability in WordPress with minimal human interaction, in 6 to 10 hours of constant iteration by itself, is insane. WordPress Core is a target that vulnerability researchers hope to find a single security issue in even once in their entire lives, because doing so is neither trivial nor easy.<br>This enhanced capability has led to a complete shift in how we see security research at Assetnote/Searchlight. We had our yearly all-hands this past week, and sitting down with all of our researchers in a room, I laid out our strategic mission for the next year. I directed them to find internet melting bugs . If we don't find them, attackers really won't be that far...

security research vulnerabilities gets offensive frontier

Related Articles