Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks | Zenodo
Skip to main
You are using an outdated browser. Please upgrade your browser to improve your experience.
Planned intervention : On Thursday, August 13th, 06:15 UTC, Zenodo will be unavailable for 3-5 minutes to perform a storage cluster upgrade.
Published August 7, 2026
| Version 1.1
Preprint
Open
Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks
Authors/Creators
Jordanov, Kosta1
Yordanov, David2
Jordanova, Yana3
Show affiliations
1.
Lenz Research
2.
Bocconi University
3.
American College of Sofia
Description
Abstract
Frontier large language models achieve comparable accuracy on public benchmarks, a parity often read as evidence that they are interchangeable as assessors of factual claims. We test that assumption directly, on claims not drawn from any public benchmark. Five frontier models were each asked to adjudicate 1,000 real-world claims submitted by users to a fact-checking platform. They were tasked with assigning every claim a verdict on a five-point scale from True to False, as well as reporting their confidence in each answer.
On the 997 claims where all five models returned a usable verdict, they fail to reach consensus on 63%, and on 23% the two furthest-apart verdicts differ by at least two verdict categories. The ordinal Krippendorff’s α of 0.77 reflects structured but far from interchangeable judgement. We find that disagreement concentrates in the intermediate verdicts, where claims resist clean adjudication: the definitive poles are unanimous about half the time, against roughly one in ten for the intermediate verdicts.
In relation to disagreement, our results show that model confidence is not a reliable indicator of whether the panel will agree. The models are highly confident almost everywhere, rating 76% of answers 9 or 10 on a 1–10 confidence scale, yet they agree with one another on confidence markedly less than on the verdicts themselves (α = 0.44). The verdict that is given to an assertion will then depend heavily on which model one consults, which matters a great deal as more people turn to such models to verify information.
Key findings:
On 63% of claims (632 / 997; 95% CI 60–66%) at least one model dissents from the panel majority, or no majority forms at all.
On 23% of claims (232 / 997; 95% CI 21–26%), the two furthest-apart verdicts differ by at least two verdict categories — an actual dispute over the claim, not a difference in calibration.
The ordinal Krippendorff's α of 0.77 , across 5 models on 997 claims, reflects structured but far from interchangeable judgement.
Disagreement concentrates in the intermediate verdicts , where claims resist clean adjudication: the definitive poles are unanimous about half the time, against roughly one in ten for the intermediate verdicts.
The models are highly confident almost everywhere , rating 76% of answers 9 or 10 on a 1–10 confidence scale, yet they agree with one another on confidence markedly less than on the verdicts themselves (α = 0.44).
HTML rendering: https://lenz.io/research/llm-disagreement/v1.1<br>Harness, corpus, and raw results: https://github.com/lenzhq/lenz-research
Files
lenz-llm-disagreement-v1.1.pdf
Files<br>(371.0 kB)
Name<br>Size
Download all
lenz-llm-disagreement-v1.1.pdf
md5:249206400af657d77393c4fc00bd4f73
110.4 kB
Preview
Download
lenz-llm-disagreement.csv
md5:19d1d1646da4c569a8ac79c5436e62d4
260.6 kB
Preview
Download
Additional details
Related works
Is identical to
Preprint:
https://lenz.io/research/llm-disagreement/v1.1
(URL)
Software
Repository URL
https://github.com/lenzhq/lenz-research
Programming language
Python
Development Status
Active
414
Views
188
Downloads
Show more details
All versions<br>This version
Views
Total views
414
105
Downloads
Total downloads
188
Data volume
Total data volume
159.6 MB<br>220.8 kB
More info on how stats are collected....
Versions
External resources
Indexed in
OpenAIRE
Communities
Keywords and subjects
Keywords
LLM evaluation
frontier model evaluation
LLM disagreement
LLM-as-judge
fact-checking
benchmark contamination
Details
DOI
DOI Badge
DOI
10.5281/zenodo.21829261
Markdown
[](https://doi.org/10.5281/zenodo.21829261)
reStructuredText
.. image:: https://zenodo.org/badge/DOI/10.5281/zenodo.21829261.svg<br>:target: https://doi.org/10.5281/zenodo.21829261
HTML
Image URL
https://zenodo.org/badge/DOI/10.5281/zenodo.21829261.svg
Target URL
https://doi.org/10.5281/zenodo.21829261
Resource type<br>Preprint
Publisher<br>Lenz Research
Languages
English
Rights
License
Creative Commons Attribution 4.0 International
The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited.
Read...