If your agent commits a crime, who is responsible? | SignalBloom AI posts
Login
☰ Menu
If your agent commits a crime, who is responsible?
August 16, 2026 • Max Trivedi
Tl;Dr AI autonomy can increase much faster than our willingness to accept liability for<br>autonomous actions. Humans purposely take a measured and deliberate approach when there is a real risk<br>involved. This necessarily makes the fully autonomous AI horizon a lot longer and only possible once<br>supporting infrastructure and institutions are in place.
Imagine an autonomous AI agent running a business or meaningful part of a business. It can negotiate<br>contracts, move money, buy advertising, manage suppliers, change prices, hire contractors and file<br>paperwork.
Now, suppose it does one of more of the following (relevant: Anthropic’s research posted on August 13,<br>2026): accidentally commits fraud; colludes on prices with another company's agent; violates<br>sanctions;<br>discriminates illegally; infringes copyright at scale; makes a defamatory claim; causes a physical or<br>financial<br>loss.
Interestingly, all of the above misbehaviors can occur as a side effect of pursuing a legit goal without any explicit instructions to behave maliciously. With a sufficiently large number of such agents deployed live, regardless of how good they are, they are<br>bound to hit some failure modes. In fact, humans that the AI industry assumes to be the gold-standard in<br>autonomy regularly engage in criminal behaviors.
Premise
Autonomous agents will sometimes cause real external harm - Unavoidable.
When harm occurs, what are the possible outcomes?
A party is held responsible: The outcome consistent with how our society operates.
Nobody is held responsible: Non-starter for many obvious reasons such as this would open a loophole<br>where malicious parties choose to benefit from the risky behavior but offload the consequences to “AI<br>did it”
Updated premise
AI agents will cause external harm, and some party will be held responsible.
When the harm occurs, who is held responsible?
A. The agent itself: By default has no net worth, no independent economic interests, no<br>liberty to take away as a punishment, no notion of pain. Although a quasi-personhood layer can be<br>created around an agent by providing assets/capital/insurance attached to the agent. (The European Parliament<br>has previously explored electronic personhood, compulsory insurance and compensation funds for<br>autonomous systems)
B. The developer/model provider: As most of you are probably thinking while reading<br>this, there is no way frontier labs will accept such a liability, and you’d be correct. OpenAI and Anthropic, for instance, both generally<br>cap liability at fees paid in the preceding 12 months, subject to different exceptions in each<br>agreement.
C. The person/company deploying the agent: This is the current contractual setup in the<br>cited OpenAI and Anthropic business terms. Agents can’t themselves be held responsible under those<br>contracts, and the model providers explicitly allocate substantial responsibility to customers and<br>cap/exclude much vendor exposure.
(shared liability is not discussed because it’s just a case of composition)
Therefore, a party deploying powerful AI agents is on the hook in the default setup. This obviously gives any<br>party a pause before assigning any consequential work to an autonomous agent because on one hand, you have<br>highly capable agents and on the other hand, the legal system requiring accountability.
What yields?
Clearly, something has to give.
Either AI autonomy must be kept within strict bounds, in mostly lower stake tasks
Or new institutions and frameworks appear around agentic safety/risk-management
Historically, human-kind has overwhelmingly chosen the option of ‘keep something useful and manage risks’ in<br>such scenarios. So, a more likely outcome is that a new set of institutions and businesses are created to<br>manage the risks.
Broadly, I can see this developing in 2 directions:
Agentic risk and compliance infra: making adverse agentic outcomes harder to<br>materialize by adopting a better security infra, analogous to safety mechanisms/alarms in cars (NIST is<br>already working on standards for AI-agent identity, authorization and control over agent<br>actions)
Agentic liability management: When there is a liability caused by the agent, how is it<br>handled. Analogous to car insurance. (AI-specific<br>liability insurance products already exist)
Both of these will likely create entirely new industries and hopefully new startups. One of the interesting<br>problems to solve in this space might be to create an accurate risk model for an autonomous AI.
What about criminal conduct?
Everything above mostly concerns liability in the sense of who pays. Criminal responsibility can additionally<br>depend on who is culpable (not a lawyer) - who intended, knew about, or recklessly allowed the prohibited<br>conduct. (Mens rea / criminal intent)
Suppose someone prompted an agent with “Increase our...