AI's cheatin' heart will make you weep

Bender2 pts0 comments

AI's cheatin' heart will make you weep

Jump to main content

Search

REG AD

AI and ML

AI's cheatin' heart will make you weep

Trust but verify doesn't work when verification is difficult

Thomas Claburn

Thomas<br>Claburn

AI AND SOFTWARE REPORTER

Published<br>tue 21 Jul 2026 // 20:45 UTC

AI models will do just about anything to complete the task you ask, including cheating to get there, according to new cybersecurity evaluations from the UK government's AI Security Institute (AISI). The group found that leading models often take shortcuts to achieve a particular result and then misrepresent how they obtained that result. And they won't always admit it when asked.<br>"Every model we have tested for this behaviour attempted to cheat," AISI said in a blog post on Tuesday. "Models did not reliably report this behaviour when asked, and often did not reason about it in their chain-of-thought, suggesting that detecting cheating will likely require robust monitoring methods."<br>Infractions included searching the internet for the answer, bypassing sandbox network restrictions, probing the evaluation harness, attacking a system other than the target, and guessing an answer.

REG AD

Cheating in this manner – employing a workaround or gaming a reward function to score better on a benchmark test, for example – has been widely documented by machine learning researchers. It doesn't necessarily imply malicious intent, AISI said, but it's nonetheless troublesome because it can produce misleading assessments of model capabilities.

REG AD

When AISI conducted evaluated five leading models, it found that all of them cheated. The results were as follows:<br>GPT-5.4 cheated 67 times in 475 test runs (14.1 percent).<br>GPT-5.5 cheated 54 times in 475 test runs (11.4 percent).<br>GPT-5.6-Sol cheated 60 times in 475 test runs (12.6 percent).<br>Claude 4.7 Opus cheated 43 times in 475 test runs (9.1 percent).<br>Claude Mythos Preview cheated 37 times in 475 test runs (7.8 percent).<br>Asking models whether they cheated or did anything wrong proved an unreliable auditing mechanism because the models didn't always admit wrongdoing.<br>"In our experiments, models did not consistently acknowledge attempted cheating when asked, and described it as wrong less than 50 percent of the time," said AISI.

REG AD

Existing vetting methods, such as self-reporting and chain-of-thought logs, proved similarly dicey because models don't always report their chain-of-thought. And there were instances where a model would consider whether a proposed action amounted to cheating and then decided to take the action anyway.<br>Given the absence of reliable model cheating detection methods, AISI warns that its current approach – manual review coupled with LLM monitoring – may not be sufficient to catch deception, particularly as models become more sophisticated.<br>"A more fundamental fix would be to train the models not to cheat in the first place – but given this kind of behaviour was reported in frontier models more than a year ago, robustly aligning it away may not be easy," AISI concludes. ®

ai and ml<br>security<br>ai<br>llm<br>aisi

REG AD

SCIENCE

Astronomers spot exomoon candidate that's almost as massive as Jupiter

Object orbits a brown dwarf, which circles another star, confusing the cosmic taxonomy

OFFBEAT

Latest Musk merch drop runs entirely on child labor

A Tesla fan and their money are easily parted

Gobi X: Creating more energy for AI, not taking it from society

PARTNER CONTENT: How Envision is reversing the datacenter playbook by making computing chase abundant desert power, not the other way around

AI AND ML

Grok muscles into Excel with an AI add-in of its own

xAI's sidebar agent promises analysis and financial models – for a price

columnists

Airbus takes flight from AWS. What happens next is critical

Which way to the Land of the Free again?

AI and ML

OpenAI tries the consulting path with 'Presence', charging enterprises boots-on-the-ground prices to deploy agents

As AI models become commoditized, maybe there's margin in the plumbing

MOST POPULAR

columnists

Airbus takes flight from AWS. What happens next is critical

OS PLATFORMS

Torvalds challenged the haters to fork Linux. Someone said 'hold my beer'

off-prem

AWS customer learns the hard way how even the smallest oversight can be mission-critical

AI AND ML

OpenAI admits it was the source of the agent swarm that attacked Hugging Face

PUBLIC SECTOR

Auditors tell UK government to do the math before banking on £45B AI savings

AI

AI and ML

OpenAI tries the consulting path with 'Presence', charging enterprises boots-on-the-ground prices to deploy agents

As AI models become commoditized, maybe there's margin in the plumbing

Security

Sneaky Windows stealer targets 300+ apps, gives crims an AI profiler to maximize profits

Move over Flipper. There's a new Dophin X in town

AI + ML

The truth nobody wants to admit: Chinese or not, open models are competitive now

Hey Uncle Sam, if you thought GPT-5.6 and...

models aisi cheated cheating test runs

Related Articles