Who decides when the task is done? Moving completion authority to checkers

kushagrab211 pts0 comments

Who Decides When theTask Is Done?: Measuring the Effect of Moving Completion Authority from LLMs to Checkers | Zenodo

Skip to main

You are using an outdated browser. Please upgrade your browser to improve your experience.

Published July 30, 2026

| Version v1

Publication

Open

Who Decides When theTask Is Done?: Measuring the Effect of Moving Completion Authority from LLMs to Checkers

Authors/Creators

Bhatnagar, Kushagra<br>(Researcher)

Description

An LLM agent finishes a task in one of two ways: either the model judges its own

work complete, which we call advisory feedback, or a checker computes completion

and the model&rsquo;s claim carries no authority, which we call binding feedback. The

distinction matters because an LLM is a generator rather than a verifier. When the

specification is incomplete, it fills the gap from its priors, grades its work against

that reconstructed answer key, and can confidently declare wrong work finished. We

call this failure the false-DONE: the model does not know when it does not know.

Across three preregistered experiments on frozen buggy-Python repair corpora, 2,862

episodes over seven models from five providers at under one dollar of total compute,

we isolate this authority placement as the only varying factor. Binding lifts a weak

model by 9.2 points while leaving a strong one unchanged. Across a seven-model

capability ladder the gain is a window, peaking mid-ladder at +14.9 points, zero

at the ceiling, and negative for the weakest model, which fails by thrashing rather

than false claiming. Raising task difficulty by composing two and three simultaneous

bugs roughly doubles the window&rsquo;s depth, upto +31.7 points, but never moves its

ceiling : the strongest model produces zero false claims at every difficulty. The window

is therefore a property of the model, not the model–task pair, and the advisory

false-DONE rate mediates the entire effect, which makes the intervention predictable:

a short probe that counts false claims tells a designer in advance whether completion

gating will pay.

Files

who_decides_when_the_task_is_done.pdf

Files<br>(940.2 kB)

Name<br>Size

Download all

who_decides_when_the_task_is_done.pdf

md5:9fd73bc66f9b59cd988bc2af50a93fd2

940.2 kB

Preview

Download

Additional details

Software

Repository URL

https://github.com/kushagrab21/binding-feedback-experiment

Programming language

Python

Views

Downloads

Show more details

All versions<br>This version

Views

Total views

Downloads

Total downloads

Data volume

Total data volume

0 Bytes<br>0 Bytes

More info on how stats are collected....

Versions

External resources

Indexed in

OpenAIRE

Communities

Details

DOI

DOI Badge

DOI

10.5281/zenodo.21698264

Markdown

[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.21698264.svg)](https://doi.org/10.5281/zenodo.21698264)

reStructuredText

.. image:: https://zenodo.org/badge/DOI/10.5281/zenodo.21698264.svg<br>:target: https://doi.org/10.5281/zenodo.21698264

HTML

Image URL

https://zenodo.org/badge/DOI/10.5281/zenodo.21698264.svg

Target URL

https://doi.org/10.5281/zenodo.21698264

Resource type<br>Publication

Publisher<br>Zenodo

Rights

License

Creative Commons Attribution 4.0 International

The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited.

Read more

Citation

Export

Technical metadata

Created

July 30, 2026

Modified

July 30, 2026

Jump up

This site uses cookies. Find out more on how we use cookies

Accept all cookies<br>Accept only essential cookies

zenodo model https done completion authority

Related Articles