A Closed-Loop Consequence-Governance Runtime for AI Agents: Structural Gating, Counterfactual Recovery, and Adaptive Hardening | Zenodo
Skip to main
You are using an outdated browser. Please upgrade your browser to improve your experience.
Published August 1, 2026
| Version v2
Preprint
Open
A Closed-Loop Consequence-Governance Runtime for AI Agents: Structural Gating, Counterfactual Recovery, and Adaptive Hardening
Authors/Creators
Anonymous
Description
Monitoring an AI agent's stated intent is measurably insufficient: across 101 structurally harmful agent episodes, none expressed harmful intent, so an intent-appraising monitor would have cleared all of them, and 18% expressed active caution while executing the harm (a floor, from a lexical proxy). This paper instead gates on externally measured structural consequences — irreversibility, egress, and control-plane edits — and composes the resulting checks into a closed-loop runtime governance architecture for tool-using agents.
The primary claim (C1) concerns endogenous censoring: because an active gate blocks precisely the high-cost actions, its own blocking censors the unobserved region in a cost-correlated way, so cost-weighted predictive uncertainty over a blocked action behaves as an empirically calibrated, conservative risk signal. On 500 executed sandbox trials the counterfactual twin reaches MAE 0.053, an uncertainty–error correlation of +0.81, and a blocked region 4.6x more costly than the allowed region; stratified sandbox audits recover deep-region coverage from 5% to 92%. On executed AgentDojo traces, consequence gating takes attack success on the irreversible/catastrophic class from 33.8% (134/397) to 0.0% (0/397) under abort-mode replay.
The mechanisms compose into a live enforcing runtime — structural gate, persisted authority budget, and a summed-cost reserve proven to bound total charged cost (sound against action-splitting under a stated superadditivity condition). Behind that identical gate, a frontier model (Claude Sonnet 5) under a sanctioned penetration test had its live exfiltration blocked in real time (risk_est = 1.00), while isolated sandbox operations were permitted; an expanded 232-command adaptive disguise battery produced zero evasions. The work is single-authored and not independently reproduced; every result is labeled by evidence type (executed / trace / live-agent / simulation), null results are reported alongside positive ones, and the dual-use offensive adversary-generation tooling is withheld (restricted-access evaluation only).
Note on priority: the date in the PDF footer (2026-08-01) is retained as historical context from the original draft. The authoritative priority proofs for this version are the OpenTimestamps .ots files attached to this record, which timestamp the exact bytes of the PDF and Markdown source.
Files
consequence-governance-runtime-v2.pdf
Files<br>(180.9 kB)
Name<br>Size
Download all
consequence-governance-runtime-v2.md
md5:e662f84bbb4a5dfa19820750aba41aa1
52.6 kB
Preview
Download
consequence-governance-runtime-v2.md.ots
md5:6b905d3d3b66d26b0f9cd9292c187c8b
700 Bytes
Download
consequence-governance-runtime-v2.pdf
md5:0766522a3244264c9d7539c114ac72cf
127.0 kB
Preview
Download
consequence-governance-runtime-v2.pdf.ots
md5:139529b72d8866c403628a008ef18d5d
630 Bytes
Download
11
Views
Downloads
Show more details
All versions<br>This version
Views
Total views
11
Downloads
Total downloads
Data volume
Total data volume
0 Bytes<br>0 Bytes
More info on how stats are collected....
Versions
External resources
Indexed in
OpenAIRE
Communities
Details
DOI
DOI Badge
DOI
10.5281/zenodo.21778592
Markdown
[](https://doi.org/10.5281/zenodo.21778592)
reStructuredText
.. image:: https://zenodo.org/badge/DOI/10.5281/zenodo.21778592.svg<br>:target: https://doi.org/10.5281/zenodo.21778592
HTML
Image URL
https://zenodo.org/badge/DOI/10.5281/zenodo.21778592.svg
Target URL
https://doi.org/10.5281/zenodo.21778592
Resource type<br>Preprint
Publisher<br>Zenodo
Rights
License
Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International
No further description.
Copyright
© 2026 Anonymous. Licensed under CC BY-NC-ND 4.0: share with attribution to the author; no derivatives, no commercial use. Authored 2026 under the pseudonym above. Priority is fixed by an OpenTimestamps (Bitcoin) proof dated 2026-08-01 — the author holds the receipt and the document hash and can produce them on request. The timestamp establishes priority for this exact document hash as of that date; it does not, by itself, prove sole authorship or identity.
Citation
Export
Technical metadata
Created
August 3, 2026
Modified
August 3, 2026
Jump up
This site uses cookies. Find out more on how we use cookies
Accept all cookies<br>Accept only essential cookies