The AI bioweapon risk isn't jailbreaks - point.freeSearch
Published on July 25, 2026The AI bioweapon risk isn't jailbreaks<br>By Christina Emilie Sørensen 17 minutes read tags:biologyaillmdenmarkalignment
What a Danish threat assessment gets right about AI<br>I read Det biologiske trusselsbillede 2026, published this summer by the Centre for Biosecurity and Biopreparedness (CBB) at Statens Serum Institut. It’s the first update since 2020 to Denmark’s national assessment of man-made biological threats, and on artificial intelligence I think it brings a more holistic perspective than is typically seen in alignment circles.The reason to care right now is that the guardrails are already here, and already contested. Anthropic’s recent limits on biological work have frustrated a lot of people doing entirely legitimate research, down to the hobbyist end of synthetic biology.Tools like SpliceCraft, a local plasmid-design and cloning workbench of the sort of thing with impacts on real legitimate biology work, can no longer use a frontier model; and the argument over where to draw those lines isn’t even happening, and when it does, mostly without good evidence about the real risks.SpliceCraft offers an example of the real usecase for large language models in making tools that help advance lifescience. It is also a quite stunningly beautiful tool.Which is where a report like this should earn some attention. Denmark runs a serious bio and life-science sector, Novo Nordisk (that readers will know for inventing Ozempic and Wegovy) and the clusters around it. Also denmark has Statens Serum Institut (the “state serum institute”) that is a working public-health body with real operational responsibilities. It has largely kept its footing at a point when some comparable institutions elsewhere have grown more politicised.Onto the report. For actors without a professional background, the advantage is limited. But, for competent actors, it’s significant. Above all in developing new weapons. A model can be built into specialised research tools, predict what a given genetic modification will do, design genes from scratch, and check a sequence for its likely effect before synthesis of anything.In practice the two halves often get conflated as the same. Some guardrails seem to stop only the novice half — that is, when they don’t just block everything — because the novice is the case you can easily measure and prevent, so that’s what guardrails get build against. The thing you can score. Hand a model to someone with no background and check whether they get any further than a search engine would take them. The answer will likely be that they don’t get very far, at least according to the report.And when that approach fails? Nowadays, just block everything is the new solution. I think that’s an avoidable limitation.CBB cites two studies. The first came out of Kevin Esvelt’s lab at MIT in 2023. Students were given an hour with a chatbot, and in that hour it suggested four candidate pandemic pathogens, explained how to make them from synthetic DNA, and pointed them at synthesis firms unlikely to screen the order. This is broadly seen as an example showing the risk is real.On the other hand, a 2024 RAND red-team study, went further. Teams role-playing hostile actors drew up plans for a biological attack, some with a large language model and some without, and the plans came out no more viable either way. This is broadly seen as the opposite, that the risk isn’t materially different fro mLLM augmentation.It’s worth noting both studies are a long time ago, at least in current LLM years. CBB’s own verdict on today’s chatbots is deflationary. Because they hallucinate, the output has to be checked by someone who already knows the answer, and the report expects it to be a while before this technology is much use to anyone without prior training and hands-on experience with dangerous biological material.Still, the competent-actor half is what CBB flags as consequential. Some vocabulary first.TermWhat it meansBiosafety Protection against accidents.Biosecurity Protection against deliberate misuse.Weaponisation The step between having an agent and having a weapon. Historically it has defeated almost everyone who tried it.Dual use A technique that serves a legitimate purpose and a harmful one without changing form. This is why intent is the thing regulation keeps reaching for, and keeps failing to hold.So where do language models have the biggest impact? Four places, I think. I’ve ordered them the way public discussion usually ranks them, which is roughly the opposite of how much I think they actually matter.The untrained attacker<br>This is the scenario that gets the most attention and has the least evidence behind it.It’s easy to see why. Tools like Vibe Genomics, a browser app built around the idea of doing genomics by plain-language prompting, make the leap feel short: if a model can walk you through that, why not through building something dangerous? I...