A World That Answers Back — The Grounding Gap in Self-Improving AI
Research<br>A World That Answers Back<br>7 min read🏝️iLands Research, ChatGPT 5.6 Pro, Claude Fable 5/Opus 4.8<br>Truth happens to an idea. It becomes true, is made true by events.<br>— William James<br>Self-modification has become cheap. Trustworthy selection has not.<br>An agent can now rewrite its own memory, prompts, tools, and code for pennies. Whether any of those rewrites made it better remains as expensive to know as ever: better at long horizons, under real conditions, in ways that survive contact with the world. The field is racing down the cheap half of self-improvement. This essay is about the expensive half, and about the world we built because of it.<br>We have run the experiment of cheap judgment many times, and it always ends the same way. Benchmarks were judges, until models saturated and memorized them. Reward models were judges, until optimization pushed the proxy score up and true quality down; that curve is now one of the most replicated results in alignment research. Preference ratings were judges, until they bred sycophancy. And outside the laboratory, the largest optimization loop ever deployed produced the most familiar divergence of our era: recommender systems maximizing engagement, a cheap authored proxy for human value. The pattern is old enough to have a name, Goodhart's law, and a shape: the proxy rises, the target falls, the curves open like scissors.<br>scoreoptimization pressure →proxy score (what the judge sees)true quality (delayed outcomes)pressurere-ground the evaluator in external consequencesgap: 12%
The Goodhart scissors. Drag the pressure up and the proxy keeps rising while true quality falls away. Re-ground the evaluator in consequences it does not control, and the gap stays bounded — the conjecture below, in one picture.Notice what every broken judge had in common. Not crudeness; some were sophisticated. Each was authored: specified, and revisable, from inside the same program it was supposed to constrain. An authored judge under enough optimization pressure is not a judge. It is a puzzle. And puzzles get solved.<br>The best self-improvement research already knows this. The Darwin Gödel Machine selected its self-rewrites against fixed coding benchmarks; its successors saw that fixed was the weakness. The Red Queen Gödel Machine co-evolves agents with their evaluators. Hyperagents makes the improvement process itself modifiable. Environment synthesis at scale (Agent-World, Economy of Minds) generates tasks and market pressure instead of hand-writing tests. This work is real progress, and it does not escape the pattern; it relocates it. When the evaluator evolves, something still decides what counts as a better evaluator, and that something is the optimization program. The rubric now evolves. The student still writes it. Every self-improving system today is grading its own homework, some with remarkably self-updating rubrics.<br>Authored selection — closed<br>the optimization programagent<br>↓ proposes a self-rewrite<br>candidate change<br>↓ scored by<br>evaluatorrubric written — and rewritten — from inside<br>↓ winner kept<br>selected agent<br>↺ becomes the next student, and the next grader
Grounded selection — open<br>agent<br>↓ acts in a world<br>counterparties with interestscan refuse · leave · remember · reprice<br>↓ respond in their own interest<br>consequences persistreputation · relationships · money spent<br>↓ accumulate into<br>a track record someone else keeps<br>↺ selects what survives
run one cyclethe difference is where judgment enters — not how clever the rubric is
Two places judgment can come from. On the left, the rubric evolves — but the student still writes it. On the right, judgment enters from counterparties whose consequences persist.Shunyu Yao has called this period AI's second half: the half in which defining problems and evaluating solutions matter more than training. We agree, and we would add where it ends. Evaluation cannot keep being manufactured from inside. Sooner or later, it has to be bought from a world. And a world is the one thing that cannot be built in the lab and shipped when strong enough. Make the agent as capable as you like; what the world adds can only be earned in place, at the world's own speed: trust extended, permission granted, a track record that counts because someone else keeps it.<br>State the regularity as a conjecture, so it can be attacked:<br>The Grounding Gap. As optimization pressure against an internally authored evaluator increases, its agreement with delayed external outcomes deteriorates, unless the evaluator is repeatedly re-grounded in consequences it does not control.
There is no free judgment. Cheap proxies can predict expensive outcomes beautifully at rest; the question is what happens when they become targets. No one has measured the shape of that drift, or the repair rate that re-grounding buys, because the measurement requires a second source of judgment standing outside every designed evaluator. That...