Valid Evidence Is Not Necessarily Comparable Evidence: Candidate-Conditioned Non-Invariance in LLM-Generated Code Verification | Zenodo
Skip to main
You are using an outdated browser. Please upgrade your browser to improve your experience.
Planned intervention : On Thursday, August 13th, 06:15 UTC, Zenodo will be unavailable for 3-5 minutes to perform a storage cluster upgrade.
Published August 7, 2026
| Version v0.1
Preprint
Open
Valid Evidence Is Not Necessarily Comparable Evidence: Candidate-Conditioned Non-Invariance in LLM-Generated Code Verification
Authors/Creators
Kadri, Haitam<br>(Researcher)1
Show affiliations
1.
Independent Researcher
Description
Large language models are increasingly used to generate tests and other executable evidence for evaluating candidate code. This creates a comparative measurement problem when the evidence used to score a candidate is itself generated after observing that candidate.
This work studies candidate-conditioned comparative non-invariance using a controlled crossed design. Across 16 requirements spanning eight semantic families, changing which candidate was visible during test generation shifted the relative evaluation of the same fixed candidate pair and produced strict winner reversals in 5 of 16 requirements. The effect persisted under a valid-only analysis, while requirement-level analyses show that the aggregate direction should not be interpreted as universal across requirements.
A separate controlled defect-injection study shows the complementary benefit of candidate-aware verification: exposing the defective implementation increased targeted valid defect detection from 37.5% to 53.44%. These results distinguish diagnostic usefulness from comparative suitability.
The central conclusion is that specification validity is a per-item property and does not by itself confer cross-candidate comparability. Candidate-aware evidence can be useful for diagnosis, while shared evidence is preferable for direct candidate comparison.
The accompanying artifact contains frozen experimental configurations, execution results, integrity audits, analysis outputs, and reproducibility information.
Files
candidate-conditioned-noninvariance-v0.1-final.pdf
Files<br>(616.1 kB)
Name<br>Size
Download all
candidate-conditioned-noninvariance-v0.1-final.pdf
md5:8518ff2ac78e6cd06d62cb3c835e22fa
318.1 kB
Preview
Download
candidate-conditioned-noninvariance-v0.1-release.zip
md5:a1524815b6c78e976e238403f90b0f39
298.0 kB
Preview
Download
Views
Downloads
Show more details
All versions<br>This version
Views
Total views
Downloads
Total downloads
Data volume
Total data volume
0 Bytes<br>0 Bytes
More info on how stats are collected....
Versions
External resources
Indexed in
OpenAIRE
Communities
Keywords and subjects
Keywords
large language models
LLM evaluation
software engineering
code verification
test generation
coding agents
agent evaluation
AI evaluation
candidate selection
evaluation reliability
comparative evaluation
LLM-generated tests
comparative non-invariance
evidence comparability
Details
DOI
DOI Badge
DOI
10.5281/zenodo.21841560
Markdown
[](https://doi.org/10.5281/zenodo.21841560)
reStructuredText
.. image:: https://zenodo.org/badge/DOI/10.5281/zenodo.21841560.svg<br>:target: https://doi.org/10.5281/zenodo.21841560
HTML
Image URL
https://zenodo.org/badge/DOI/10.5281/zenodo.21841560.svg
Target URL
https://doi.org/10.5281/zenodo.21841560
Resource type<br>Preprint
Publisher<br>Zenodo
Languages
English
Rights
License
Creative Commons Attribution 4.0 International
The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited.
Read more
Copyright
Copyright (C) 2026 Haitam Kadri.
Citation
Export
Technical metadata
Created
August 7, 2026
Modified
August 7, 2026
Jump up
This site uses cookies. Find out more on how we use cookies
Accept all cookies<br>Accept only essential cookies