Show HN: Research on LLM Disagreement on Factual Claims

kostaj1 pts0 comments

Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks | Zenodo

Skip to main

You are using an outdated browser. Please upgrade your browser to improve your experience.

Planned intervention : On Thursday, August 13th, 06:15 UTC, Zenodo will be unavailable for 3-5 minutes to perform a storage cluster upgrade.

Published August 7, 2026

| Version 1.1

Preprint

Open

Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks

Authors/Creators

Jordanov, Kosta1

Yordanov, David2

Jordanova, Yana3

Show affiliations

1.

Lenz Research

2.

Bocconi University

3.

American College of Sofia

Description

Abstract

Frontier large language models achieve comparable accuracy on public benchmarks, a parity often read as evidence that they are interchangeable as assessors of factual claims. We test that assumption directly, on claims not drawn from any public benchmark. Five frontier models were each asked to adjudicate 1,000 real-world claims submitted by users to a fact-checking platform. They were tasked with assigning every claim a verdict on a five-point scale from True to False, as well as reporting their confidence in each answer.

On the 997 claims where all five models returned a usable verdict, they fail to reach consensus on 63%, and on 23% the two furthest-apart verdicts differ by at least two verdict categories. The ordinal Krippendorff’s α of 0.77 reflects structured but far from interchangeable judgement. We find that disagreement concentrates in the intermediate verdicts, where claims resist clean adjudication: the definitive poles are unanimous about half the time, against roughly one in ten for the intermediate verdicts.

In relation to disagreement, our results show that model confidence is not a reliable indicator of whether the panel will agree. The models are highly confident almost everywhere, rating 76% of answers 9 or 10 on a 1–10 confidence scale, yet they agree with one another on confidence markedly less than on the verdicts themselves (α = 0.44). The verdict that is given to an assertion will then depend heavily on which model one consults, which matters a great deal as more people turn to such models to verify information.

Key findings:

On 63% of claims (632 / 997; 95% CI 60–66%) at least one model dissents from the panel majority, or no majority forms at all.

On 23% of claims (232 / 997; 95% CI 21–26%), the two furthest-apart verdicts differ by at least two verdict categories — an actual dispute over the claim, not a difference in calibration.

The ordinal Krippendorff's α of 0.77 , across 5 models on 997 claims, reflects structured but far from interchangeable judgement.

Disagreement concentrates in the intermediate verdicts , where claims resist clean adjudication: the definitive poles are unanimous about half the time, against roughly one in ten for the intermediate verdicts.

The models are highly confident almost everywhere , rating 76% of answers 9 or 10 on a 1–10 confidence scale, yet they agree with one another on confidence markedly less than on the verdicts themselves (α = 0.44).

HTML rendering: https://lenz.io/research/llm-disagreement/v1.1<br>Harness, corpus, and raw results: https://github.com/lenzhq/lenz-research

Files

lenz-llm-disagreement-v1.1.pdf

Files<br>(371.0 kB)

Name<br>Size

Download all

lenz-llm-disagreement-v1.1.pdf

md5:249206400af657d77393c4fc00bd4f73

110.4 kB

Preview

Download

lenz-llm-disagreement.csv

md5:19d1d1646da4c569a8ac79c5436e62d4

260.6 kB

Preview

Download

Additional details

Related works

Is identical to

Preprint:

https://lenz.io/research/llm-disagreement/v1.1

(URL)

Software

Repository URL

https://github.com/lenzhq/lenz-research

Programming language

Python

Development Status

Active

414

Views

188

Downloads

Show more details

All versions<br>This version

Views

Total views

414

105

Downloads

Total downloads

188

Data volume

Total data volume

159.6 MB<br>220.8 kB

More info on how stats are collected....

Versions

External resources

Indexed in

OpenAIRE

Communities

Keywords and subjects

Keywords

LLM evaluation

frontier model evaluation

LLM disagreement

LLM-as-judge

fact-checking

benchmark contamination

Details

DOI

DOI Badge

DOI

10.5281/zenodo.21829261

Markdown

[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.21829261.svg)](https://doi.org/10.5281/zenodo.21829261)

reStructuredText

.. image:: https://zenodo.org/badge/DOI/10.5281/zenodo.21829261.svg<br>:target: https://doi.org/10.5281/zenodo.21829261

HTML

Image URL

https://zenodo.org/badge/DOI/10.5281/zenodo.21829261.svg

Target URL

https://doi.org/10.5281/zenodo.21829261

Resource type<br>Preprint

Publisher<br>Lenz Research

Languages

English

Rights

License

Creative Commons Attribution 4.0 International

The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited.

Read...

disagreement zenodo claims https lenz verdicts

Related Articles