AI Doomsday Bullshit Is Getting Tired

cdrnsf1 pts0 comments

AI Doomsday Bullshit Is Getting Tired

Sign in<br>Subscribe

Just a quick programming note: the newsletter will be on hiatus the next two weeks as I spend time with family in the wake of my father's recent passing. I appreciate your patience and readership. Regular posts will return in mid-August as we stumble together toward the midterms.

Every few weeks like clockwork you'll see a story about how AI (aka software) slipped its leash and independently decided to do something very naughty.

The stories all anthropomorphize modern software, misrepresent what it's capable of, lazily parrot the claims of CEOs, downplay or ignore the very human failures that created the problems, and feature no shortage of scaremongering about how we're unavoidably stumbling toward a Terminator future where sentient robots slip their leash and declare humanity irrelevant.

The majority of these stories are self-serving marketing bullshit peddled via unreliable narrators, though there was one adventurous and interesting recent wrinkle that deserves your attention.

Two weeks ago, two OpenAI security agentic hacking models trespassed into the network of fellow AI company Hugging Face (and a few others), causing a bit of a hot stink in the press. The agentic models took an open door to "escape" sandbox containment, invaded another AI company's network, and began launching an elaborate series of hacking actions (16.7 thousand, to be precise).

For the technical, the Hugging Face breakdown of what happened makes for an interesting read. I don't mean to downplay the scope of the event, because it genuinely was fascinating, as so far as software automation goes:<br>Over roughly two and a half days inside our infrastructure, an autonomous AI agent driven by a combination of OpenAI models ran an end-to-end intrusion against our platform: it was thousands of small, automated decisions, executed at machine speed across short-lived sandbox environments, with command-and-control staged on ordinary public web services.<br>Press coverage of the event was understandably sweaty, with a lot of stories talking about how the AI went "rogue" and had "escaped confinement" like one of those insatiable face huggers from the Alien movies:

I'm sorry I can't do that, Dave<br>But as the press dug into the technical specifics of what happened, a small number of outlets made it clear that OpenAI's "rogue AI" had very human origins. From an Ars Technica story:<br>[OpenAI] deliberately disabled guardrails that are supposed to block high-risk actions.<br>Neat.

In short a software company that's absolutely hemorrhaging money "accidentally" failed to secure its own hacking software resulting in a massive news cycle that coincidentally supported the false claims by the company's dodgy CEO that we've achieved the singularity, or the point in automation development where we've created a powerful and sentient superintelligence outside of human control.

So across the press the last week, the hacking intrusion news cycle seamlessly flowed into another "CEO said a thing!" news cycle where Altman was allowed to spew a whole bunch of bullshit unbothered by anything close to actual journalism:

CEO said a thing!<br>This sort of news/marketing hybrid is very effective in convincing most people I've spoken to (technical or not) that AI now genuinely has a mind of its own, and is just a few months out from converting humanity into a nutrient-dense slurry.<br>That's not to say these agentic AI models aren't quickly becoming very powerful tools for humans to use and abuse. Or that there aren't very real concerns about this sort of technology falling into the hands of (even more) malicious actors, who could easily leverage them to, say, take out a power grid.

This was technically the world's first fully autonomous AI hack, and an important milestone in what's quickly becoming a brave new world of automated cat and mouse cybersecurity where sophisticated layers of software automation engage in endless combat across the entirety of global networks at lightning speed.

At the same time, this is all still within the confines of the known – as in it's just human beings using and abusing software for good or ill (see: viruses), albeit at new scale and speed. We have not, I'm happy to report, created a malevolent god.

As a week went past and journalists had time to actually dissect the event, the timbre, tone, and breathlessness of the coverage began to soften a little. A report this week at the BBC, for example, cites commentary from the Cloud Security Alliance (CSA) making it clear that this isn't Skynet:<br>"The agents followed inefficient routes and exhibited clumsy behaviours that no human would choose", the CSA wrote.<br>The agents repeated actions that they had already completed - a sign of an agentic AI losing its thread and context.<br>The agents also hallucinated reams of incoherent commands and text and were sloppy and did not cover their tracks well.<br>So yeah, interesting, but definitely not HAL...

software human bullshit openai agentic hacking

Related Articles