Language Models Agree with Each Other, Not with Readers

sbulaev1 pts0 comments

[2607.29274] Language Models Agree With Each Other, Not With Readers

Skip to main content

Search arXiv

Press Enter to search · Advanced search

-->

Computer Science > Information Retrieval

arXiv:2607.29274 (cs)

[Submitted on 31 Jul 2026]

Title:Language Models Agree With Each Other, Not With Readers

Authors:Kazuki Nakayashiki, Keisuke Watanabe<br>View a PDF of the paper titled Language Models Agree With Each Other, Not With Readers, by Kazuki Nakayashiki and 1 other authors

View PDF<br>HTML (experimental)

Abstract:Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact of the design: a crowdworker given the model's instruction is running the model's prompt. We measure convergence against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others' marks is off by default.

Agreement is the overlap between two size-matched sentence sets minus the overlap expected when each is resampled within its own depth-and-length bands. The null's calibration is demonstrated, not asserted: every pair involving a random baseline lands within 0.006 of zero. On the median document each party names 14 sentences of 70; two readers share 4.1 and two models 8.7.

Across 18 model arms spanning 11 vendors, 3 countries and both weight regimes, the median of 153 model pairs is +0.093 against a human yardstick of +0.040, and 99 sit entirely above the human interval. Two frontier models from rival labs reach +0.203, twice what GPT-4o agrees with itself on a second call. The effect is not determinism, prompt wording, procedure, vendor or routing, and it is graded: the smallest models agree at the human level. No model agrees with readers detectably more than a reader does, and at equal depth and length no surface feature separates their choices.

The multiples are procedure-dependent and the ordering is not: models are cut to their sharpest set while a reader's is a random draw from what they marked, and blunting the models alike halves the gap without closing it. Tested out of sample on four models released after this analysis, against predictions fixed beforehand, none clears the human interval. A population simulated from several models is not several populations.

Comments:<br>18 pages. Ancillary files include all three pre-registrations, every analysis script and every result artifact; the paper contains no numeric literal for a measured value and this http URL regenerates all of them from the shipped artifacts alone

Subjects:

Information Retrieval (cs.IR); Computation and Language (cs.CL); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)

Cite as:<br>arXiv:2607.29274 [cs.IR]

(or<br>arXiv:2607.29274v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2607.29274

Focus to learn more

arXiv-issued DOI via DataCite (pending registration)

Submission history<br>From: Kazuki Nakayashiki [view email]<br>[v1]<br>Fri, 31 Jul 2026 10:44:10 UTC (217 KB)

Full-text links:<br>Access Paper:

View a PDF of the paper titled Language Models Agree With Each Other, Not With Readers, by Kazuki Nakayashiki and 1 other authors<br>View PDF<br>HTML (experimental)<br>TeX Source

view license

Ancillary-file links:<br>Ancillary files (details):

HARNESS-DEFECTS.md<br>PILOT-PREREG.md<br>PREREG-PANEL.md<br>PREREG.md<br>RESULTS-PAPER7.md

algorithm-control.json<br>algorithm-control.py<br>audit-r1-truncation.json<br>audit-r1-truncation.py<br>audit-r10-method-dependence.json<br>audit-r10-method-dependence.py<br>audit-r12-unit-of-choice.json<br>audit-r12-unit-of-choice.py<br>audit-r15-ceiling.json<br>audit-r15-ceiling.py<br>audit-r18-interpretation.json<br>audit-r18-interpretation.py<br>audit-r20-literals-and-provenance.json<br>audit-r20-literals-and-provenance.py<br>audit-r21-prereg-conformance.json<br>audit-r21-prereg-conformance.py<br>audit-r22-numerical-convergence.json<br>audit-r22-numerical-convergence.py<br>audit-r23-procedure-symmetry.json<br>audit-r23-procedure-symmetry.py<br>audit-r24-regeneration.json<br>audit-r24-regeneration.py<br>audit-r26-mirror-choices.json<br>audit-r26-mirror-choices.py<br>audit-r27-ratio-and-asymmetry.json<br>audit-r27-ratio-and-asymmetry.py<br>audit-r28-final-consistency.json<br>audit-r28-final-consistency.py<br>audit-r29-dev4.json<br>audit-r29-dev4.py<br>audit-r3-label-size.json<br>audit-r3-label-size.py<br>audit-r5-truncation-noise.json<br>audit-r5-truncation-noise.py<br>audit-r7-excess-closed-form.json<br>audit-r7-excess-closed-form.py<br>audit-r8-null-size.json<br>audit-r8-null-size.py<br>audit-r9-zero-base.json<br>audit-r9-zero-base.py<br>convergence-character.json<br>convergence-character.py<br>curve.json<br>curve.py<br>degeneracy.json<br>degeneracy.py<br>dev4-out-of-sample.json<br>dev4-out-of-sample.py<br>generation-trend.json<br>generation-trend.py<br>list-models.py<br>make-results.py<br>panel-agreement.json<br>panel-agreement.py<br>panel-coverage.json<br>panel.py<br>pilot.json<br>pilot.py<br>position-control.py<br>typicality.json<br>typicality.py<br>v2tests.py

(62 additional...

audit json models human language readers

Related Articles