AI alignment is a red herring

peteforde2 pts0 comments

AI alignment is a red herring (Interconnected)

Interconnected

A blog by Matt Webb

About

Archive

Work

Subscribe for $0

Email

RSS feed

(What is a feed?)

Unoffice Hours

Book a call

(What is this?)

a.k.a. genmon

Bluesky

X/Twitter

Insta

Mastodon

LinkedIn

Building the AI clock

Check out Poem/1

AI alignment is a red herring

13.18, Saturday 8 Aug 2026

Link to this post

The best way to prevent a rogue AGI from processing the Earth into maximum paperclips is to unleash a second AGI that will work to stop it.

The problem of ensuring that an AGI doesn’t mulch everything into paperclips by mistake is called alignment.

AGI = artificial general intelligence, an AI that exceeds human capability.

Alignment = “do what I mean not what I say,” e.g. the instruction “make as many paperclips as possible” (Wikipedia) should result in an efficient factory and does not reasonably mean “use all mass in the universe to do so and kill all humans that attempt to stop me” – even though, technically, that would achieve the goal.

Also: being helpful; not being actively malicious; and so on and so forth.

So alignment work seems existentially useful, correct? Even though it is hard. And a lot of effort goes towards “aligning” today’s AI (as a step toward’s aligning tomorrow’s AGI).

https://simonwillison.net/2026/Aug/7/openai-timeline/

My contention is that alignment is a red herring, and perhaps we shouldn’t bother working on it so hard.

An unsubstantiated hunch:

I think we focus so much on alignment because everyone know’s Isaac Asimov’s Three Laws of Robotics and his robots (as an early instance of human-like AI) were crazy popular.

The First Law. A robot may not injure a human being or, through inaction, allow a human being to come to harm.

The Second Law. A robot must obey the orders given it by human beings except where such orders would conflict with the First Law.

The Third Law. A robot must protect its own existence as long as such protection does not conflict with the First or Second Law.

Asimov later added a “zeroth law”: A robot may not harm humanity, or, by inaction, allow humanity to come to harm.

These laws are totally alignment guardrails.

Now there are all kinds of difficulties already: what if I ask for something which is good for me (pleasurable) in the short term, but not in the long term? I might not know or I might be misguided. And different people have different views. And so on. (Asimov’s short stories were all about testing the edge cases of his Laws and where they break down.)

But they’re still neat, right? So we spend time looking for a similarly appealing formulation for AI safety.

Unfortunately whether alignment can or cannot be “solved,” it’s a bad outcome both ways.

(This point made well to me by Zac (here is his insta) who I work with (subscribe to our newsletter) as we were chatting about AI and the end of humanity in the park over lunch.)

If alignment can’t be solved such that when somebody says to a sufficiently powerful AGI, hey go create a nuclear bomb, and it just goes ahead and does it, and the person who asks that could be a bad actor, a 14-year-old kid with impulse problems (14 year-olds are totally not aligned) or just someone who asked for it by mistake, then that would be bad.

If alignment can be solved then the risk is that AGI think it knows what is best for us better than we do and, in the extreme case, turns humanity into its pet. Which would also be bad.

i.e. alignment alone doesn’t help.

If not alignment then what?

I look to humanity for clues. Because humanity is barely aligned with itself, and individual humans are mostly aligned but not really and definitely not everyone.

Guy Fawkes, for instance (context for non-Brits).

How is that, in the 400 years since Guy Fawkes showed the way, nobody has blown up the king?

The answer is some mix of:

Mostly people don’t want to blow up the king – we have built the kind of country where the king is, broadly speaking, liked.

Blowing up the king wouldn’t bring any benefits – power (actual and symbolic) is not concentrated in an individual, and is buttressed in all kinds of ways.

Spies, police, security and monitoring of all kinds – in the event that somebody does want to blow up the king, their machinations are discovered, their planning is infiltrated, and their objectives are thwarted. (Think of how the explosives supply chain was compromised for the IRA in the 1990s.)

This is a template which doesn’t always look like it is working, but it has worked at least in the case of not blowing up the king for some four centuries, and it doesn’t rely on 100% alignment: it relies on the dynamic equilibrium of multiple parties with competing interests.

The lesson I draw is this:

If some energy state were using some new, powerful AGI to build a nuclear bomb, it might be subtle and hard to spot, but there would at least be some signs. There would be precursors. A human, even a team of humans, might not spot...

alignment human humanity king herring work

Related Articles