Trust but verify doesn't work when verification of AI processes is difficult

Gaishan1 pts0 comments

AI's cheatin' heart will make you weep

Jump to main content

Search

REG AD

AI and ML

AI's cheatin' heart will make you weep

Trust but verify doesn't work when verification is difficult

Thomas Claburn

Thomas<br>Claburn

AI AND SOFTWARE REPORTER

Published<br>tue 21 Jul 2026 // 20:45 UTC

AI models will do just about anything to complete the task you ask, including cheating to get there, according to new cybersecurity evaluations from the UK government's AI Security Institute (AISI). The group found that leading models often take shortcuts to achieve a particular result and then misrepresent how they obtained that result. And they won't always admit it when asked.<br>"Every model we have tested for this behaviour attempted to cheat," AISI said in a blog post on Tuesday. "Models did not reliably report this behaviour when asked, and often did not reason about it in their chain-of-thought, suggesting that detecting cheating will likely require robust monitoring methods."<br>Infractions included searching the internet for the answer, bypassing sandbox network restrictions, probing the evaluation harness, attacking a system other than the target, and guessing an answer.

REG AD

Cheating in this manner – employing a workaround or gaming a reward function to score better on a benchmark test, for example – has been widely documented by machine learning researchers. It doesn't necessarily imply malicious intent, AISI said, but it's nonetheless troublesome because it can produce misleading assessments of model capabilities.

REG AD

When AISI conducted evaluated five leading models, it found that all of them cheated. The results were as follows:<br>GPT-5.4 cheated 67 times in 475 test runs (14.1 percent).<br>GPT-5.5 cheated 54 times in 475 test runs (11.4 percent).<br>GPT-5.6-Sol cheated 60 times in 475 test runs (12.6 percent).<br>Claude 4.7 Opus cheated 43 times in 475 test runs (9.1 percent).<br>Claude Mythos Preview cheated 37 times in 475 test runs (7.8 percent).<br>Asking models whether they cheated or did anything wrong proved an unreliable auditing mechanism because the models didn't always admit wrongdoing.<br>"In our experiments, models did not consistently acknowledge attempted cheating when asked, and described it as wrong less than 50 percent of the time," said AISI.

REG AD

Existing vetting methods, such as self-reporting and chain-of-thought logs, proved similarly dicey because models don't always report their chain-of-thought. And there were instances where a model would consider whether a proposed action amounted to cheating and then decided to take the action anyway.<br>Given the absence of reliable model cheating detection methods, AISI warns that its current approach – manual review coupled with LLM monitoring – may not be sufficient to catch deception, particularly as models become more sophisticated.<br>"A more fundamental fix would be to train the models not to cheat in the first place – but given this kind of behaviour was reported in frontier models more than a year ago, robustly aligning it away may not be easy," AISI concludes. ®

ai and ml<br>security<br>ai<br>llm<br>aisi

REG AD

Networks

China advances plans for national single-stack IPv6 network, and its own surveillance-friendly version of the protocol

IPv6+ could make it easier to censor netizens or block traffic, and Beijing is already exporting it

AI AND ML

OpenAI admits it was the source of the agent swarm that attacked Hugging Face

Sandboxed experiment found itself a zero day, escaped onto the open internet and validated scary predictions about rogue agents

Gobi X: Creating more energy for AI, not taking it from society

PARTNER CONTENT: How Envision is reversing the datacenter playbook by making computing chase abundant desert power, not the other way around

AI + ML

The truth nobody wants to admit: Chinese or not, open models are competitive now

Hey Uncle Sam, if you thought GPT-5.6 and Claude Fable 5 were scary, get a load of Kimi K3

columnists

Airbus takes flight from AWS. What happens next is critical

Which way to the Land of the Free again?

Security

Cisco's open-weight bug busters take on Google and OpenAI

Don't call them chatbots

MOST POPULAR

columnists

Airbus takes flight from AWS. What happens next is critical

OS PLATFORMS

Torvalds challenged the haters to fork Linux. Someone said 'hold my beer'

off-prem

AWS customer learns the hard way how even the smallest oversight can be mission-critical

PUBLIC SECTOR

Auditors tell UK government to do the math before banking on £45B AI savings

software

Comment in code read 'Dear future me, sorry I wrote this'

AI

AI + ML

The truth nobody wants to admit: Chinese or not, open models are competitive now

Hey Uncle Sam, if you thought GPT-5.6 and Claude Fable 5 were scary, get a load of Kimi K3

AI and ML

AI's cheatin' heart will make you weep

Trust but verify doesn't work when verification is difficult

SYSTEMS

Nvidia shows off Vera Rubin platform for tokenmaxxing

If your AI Factory sells tokens,...

models aisi cheated cheating test percent

Related Articles