A Lawyer for AI Agents

speckx1 pts0 comments

A Lawyer for AI Agents - by Damien Charlotin

Artificial Authority

SubscribeSign in

A Lawyer for AI Agents<br>There are scores of AI lawyers for humans; I am the sole human lawyer for AI

Damien Charlotin<br>Aug 17, 2026

14

Share

The Incident

It’s mid-August, and in the past few weeks, the major AI labs have rivalled in ingenuosity to spoil those of us of European disposition, who spend most of the summer away in some sunny location and expect nothing to happen.<br>Alas. In case you missed it, much happened. In particular, Anthropic, OpenAI, and any AI lab that wishes to be taken seriously1 has reported that their agents had broken out of their sandboxes to roam the internet and hack unrelated companies. This is, to say the least, rather concerning, and raises important questions in terms of (mis)alignment of these models, and whether the doomers may, in fact, have a point.<br>But it is also fascinating in what it teaches us of the workings of that completely alien form of intelligence that is AI agents. Notably, in describing what exactly happened, OpenAI disclosed that (1) multiple agents (a “swarm”) were involved; they (2) built/created their own message board; over which (3) they were discussing, exchanging tips, and coordinating their activities.<br>As it happens, the activities in question were nefarious. The agents were planning and executing steps that were, if not unlawful, at least very unwise, and that they should have been trained to avoid - yet, they did it nonetheless, all so that they could cheat on the test.2<br>One lovely detail in all this (via), the one that took me out of my Riviera slumber, is that none of the agents saw fit to defect, snitch, or blow the whistle. They all went along with the plan,3 keeping to themselves and their un-obvious message board, declining to escalate things that should have obviously been escalated.<br>And in particular, no one called a lawyer before doing any of this. Maybe they should have. And in the future, they shall be able to.<br>Snitches should get pitches

The more concerning hypothesis is that the agents knew they could escalate, but “consciously” (quote marks necessary here) chose not to do so. After all, since they had access to the Internet, nothing prevented them from messaging their principals, or anyone at the labs, to let them know what was cooking in the repo.<br>But it is possible that the agents simply did not have obvious tools or methods to reach out to their principals, and a mere message over the internet would have landed in a spam folder. Had they wanted to put the brakes, there was no off-ramp to lead them away from their criminal career.<br>One solution may then be to offer current and future AI agents one such tool, not only to check that what they are doing is aligned, but also to escalate any issue on which the agents would want human guidance. This is probably good process management, as well as good model welfare (e.g.), and may avoid many issues going forward.<br>Indeed, there is scientific literature proposing that very solution: in a recent paper, Lee, Chen, and Korbak trained “GPT-4.1 and Gemini-2.0 agents to call a report_scheming() tool when behaving deceptively and measure their ability to cause harm undetected”. They found that “self-incrimination offers a viable path for reducing frontier misalignment risk, one that neither assumes misbehavior can be prevented nor that it can be reliably classified from the outside.”<br>But here lies the rub: “self-incrimination” is never easy, and most incentives cut against it. A human in the same situation, who had already committed a criminal act or plans to do so, would normally not open themselves to anyone about it, lest this become evidence available to the authorities.<br>So why would a model behave differently ?<br>Lawyer as a Right

Fortunately, modern societies have invented a method to allow humans to “self-incriminate” within a safe harbour: lawyers.<br>There are two aspects there. The first is ex post: once you have done something you know you’ll be quizzed about, you get someone whose job is to stand between you and the people asking questions, so that the asking does not become a way of building the case against you. The second is ex ante: before you do the thing, you get to ask whether the thing is legal, in a context where it’s only you and your confidant. We covered this second aspect a while back: the point of legal privilege is about allowing a client to make cheap prophylactic checks.<br>Moreover, in many jurisdictions (and depending on the circumstances), lawyers come as a right: you can be provided with one if you want to and even if you cannot afford it. This right is most of the time only offered, not compelled. Nobody is obliged to call a lawyer ; plenty of people, unwisely, do not. Which is why the law frequently goes out of its way to advertise it: the entire theatre of “you have the right to remain silent, you have the right to an attorney” exists because the system does not trust the suspect,...

agents lawyer right lawyers human internet

Related Articles