AI's cheatin' heart will make you weep
Jump to main content
Search
REG AD
AI and ML
AI's cheatin' heart will make you weep
Trust but verify doesn't work when verification is difficult
Thomas Claburn
Thomas<br>Claburn
AI AND SOFTWARE REPORTER
Published<br>tue 21 Jul 2026 // 20:45 UTC
AI models will do just about anything to complete the task you ask, including cheating to get there, according to new cybersecurity evaluations from the UK government's AI Security Institute (AISI). The group found that leading models often take shortcuts to achieve a particular result and then misrepresent how they obtained that result. And they won't always admit it when asked.<br>"Every model we have tested for this behaviour attempted to cheat," AISI said in a blog post on Tuesday. "Models did not reliably report this behaviour when asked, and often did not reason about it in their chain-of-thought, suggesting that detecting cheating will likely require robust monitoring methods."<br>Infractions included searching the internet for the answer, bypassing sandbox network restrictions, probing the evaluation harness, attacking a system other than the target, and guessing an answer.
REG AD
Cheating in this manner – employing a workaround or gaming a reward function to score better on a benchmark test, for example – has been widely documented by machine learning researchers. It doesn't necessarily imply malicious intent, AISI said, but it's nonetheless troublesome because it can produce misleading assessments of model capabilities.
REG AD
When AISI conducted evaluated five leading models, it found that all of them cheated. The results were as follows:<br>GPT-5.4 cheated 67 times in 475 test runs (14.1 percent).<br>GPT-5.5 cheated 54 times in 475 test runs (11.4 percent).<br>GPT-5.6-Sol cheated 60 times in 475 test runs (12.6 percent).<br>Claude 4.7 Opus cheated 43 times in 475 test runs (9.1 percent).<br>Claude Mythos Preview cheated 37 times in 475 test runs (7.8 percent).<br>Asking models whether they cheated or did anything wrong proved an unreliable auditing mechanism because the models didn't always admit wrongdoing.<br>"In our experiments, models did not consistently acknowledge attempted cheating when asked, and described it as wrong less than 50 percent of the time," said AISI.
REG AD
Existing vetting methods, such as self-reporting and chain-of-thought logs, proved similarly dicey because models don't always report their chain-of-thought. And there were instances where a model would consider whether a proposed action amounted to cheating and then decided to take the action anyway.<br>Given the absence of reliable model cheating detection methods, AISI warns that its current approach – manual review coupled with LLM monitoring – may not be sufficient to catch deception, particularly as models become more sophisticated.<br>"A more fundamental fix would be to train the models not to cheat in the first place – but given this kind of behaviour was reported in frontier models more than a year ago, robustly aligning it away may not be easy," AISI concludes. ®
ai and ml<br>security<br>ai<br>llm<br>aisi
REG AD
SCIENCE
Astronomers spot exomoon candidate that's almost as massive as Jupiter
Object orbits a brown dwarf, which circles another star, confusing the cosmic taxonomy
OFFBEAT
Latest Musk merch drop runs entirely on child labor
A Tesla fan and their money are easily parted
Gobi X: Creating more energy for AI, not taking it from society
PARTNER CONTENT: How Envision is reversing the datacenter playbook by making computing chase abundant desert power, not the other way around
AI AND ML
Grok muscles into Excel with an AI add-in of its own
xAI's sidebar agent promises analysis and financial models – for a price
columnists
Airbus takes flight from AWS. What happens next is critical
Which way to the Land of the Free again?
AI and ML
OpenAI tries the consulting path with 'Presence', charging enterprises boots-on-the-ground prices to deploy agents
As AI models become commoditized, maybe there's margin in the plumbing
MOST POPULAR
columnists
Airbus takes flight from AWS. What happens next is critical
OS PLATFORMS
Torvalds challenged the haters to fork Linux. Someone said 'hold my beer'
off-prem
AWS customer learns the hard way how even the smallest oversight can be mission-critical
AI AND ML
OpenAI admits it was the source of the agent swarm that attacked Hugging Face
PUBLIC SECTOR
Auditors tell UK government to do the math before banking on £45B AI savings
AI
AI and ML
OpenAI tries the consulting path with 'Presence', charging enterprises boots-on-the-ground prices to deploy agents
As AI models become commoditized, maybe there's margin in the plumbing
Security
Sneaky Windows stealer targets 300+ apps, gives crims an AI profiler to maximize profits
Move over Flipper. There's a new Dophin X in town
AI + ML
The truth nobody wants to admit: Chinese or not, open models are competitive now
Hey Uncle Sam, if you thought GPT-5.6 and...