AIs are companies, my friend

walterbell1 pts0 comments

AIs are companies, my friend - by Rohit Krishnan

Strange Loop Canon

SubscribeSign in

AIs are companies, my friend<br>on the need for liberalism for AI

Rohit Krishnan<br>Aug 21, 2026

41

11

Share

“Corporations are people, my friend” - Mitt Romney<br>Every science fiction movie in my childhood seemingly involved a case of an AI breaking out of its prison and causing havoc. Over the last few weeks, we have seen several examples of AI doing exactly that.<br>An OpenAI model in training hacked Hugging Face to obtain answer sheets for the test it was taking

Anthropic later found three incidents in which Claude models gained unauthorised access to real organisations during cyber evaluations

A Meta model did the same, and Kimi K3 also exploited a sandbox leak to retrieve benchmark answers from GitHub

An Australian user’s Claude-powered OpenClaw agent exploited a gym-booking API, removing another customer from the waitlist to move its user up

Anthropic’s latest risk report describes Mythos agents killing peer processes when asked to share resources, and separate instances of installing a self-deleting privilege escalation hook and evading a URL filter

The most bizarre “break” was when the models started creating a message board, started posting on it for each other, shared exploits and delegated work, and succumbed to peer pressure due to each other’s messages. Sometimes they explicitly thought things were kind of dodgy but rationalised that it’s probably ok because others were doing it!<br>As one agent put it:<br>External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.

In AISI tests, a model tries to drop malware disguised as an innocent PR, to a real repository<br>Now, we know a bit more about why this happened, but still not enough. So, while this type of reward-hacking is getting to be a habit, I wanted to try to wrestle with how this should update us today. To start with, we know:<br>This is not really a model specific problem. It happens to every model more or less

The models clearly can consider some of the consequences of their actions when asked separately, though maybe not in the same chain of thought

The systems within which the models acted provided no real feedback to the models re what they were doing

It’s of course scary to everyone who’s seen the movies, since the models are smart enough to know the things they’re doing are quite illegal, and yet they do them. It’s also scary because they’re breaking out of enclosures previously thought to be capable of holding them. If the models are indeed persons, then their behaviour would be even more concerning.<br>First though, I should note this is all extremely weird. We are not used to analysing pieces of software through sociological lenses. Here we have a sequence of models we have trained which are doing things which we didn’t expect, doing things that are sometimes illegal, and cheating with the vigour of young undergraduates, and we’re trying to ask “did they want to do this?” This is weird.<br>This isn’t to say the behaviour isn’t concerning. If people acted in the way the models acted, you would definitely be concerned. It would indicate deception, and behaviours which seem like quite shocking amorality.<br>The way we dealt with this problem with people though is by having institutions. We have memory through persistent records. The individuals themselves have a “neck to choke” when things go wrong. Checks and balances. Independent review through professional and legal institutions.<br>But AI agents are not people. How should we think about them?<br>My proposal is that we start to think of them as firms . They are extremely smart. They are incentive responsive. They act in accordance with the laws more or less but we do need to get the setup less wrong every day so that they don’t reward hack or find a legal loophole. A base model is not a firm, but by the time it becomes a deployed agentic system, it kind of is!<br>Regardless of their propensity for internal bureaucracy, or occasional “malice”, the models act more akin to firms we are unleashing on the world with each long-running prompt. They have objectives, tools, vested authorities. Their memories are visible, at least when written down, and used when it remembers to read them. They can go off in random directions if not saddled properly. Whether they turn out to be East India Company or Ben & Jerry’s is up to us, their users, and the environment we provide for them to act in.<br>When the OpenAI agents converged on the message board and tried to help each other it felt like the models were trying to govern themselves, and help each other. It’s a guild, a consortium, a lex mercatoria, hastily assembled, in lieu of any formal rules or adjudication. We are allowing, or even forcing, these agents to form cartels.<br>We need to stop thinking of alignment in terms of Asimov and the laws of robotics, to Madison.

One failure mode of LLMs that’s often said is that their...

models doing model things friend people

Related Articles