LLM-generated tests can change which implementation wins

haitamk1 pts0 comments

Valid Evidence Is Not Necessarily Comparable Evidence: Candidate-Conditioned Non-Invariance in LLM-Generated Code Verification | Zenodo

Skip to main

You are using an outdated browser. Please upgrade your browser to improve your experience.

Planned intervention : On Thursday, August 13th, 06:15 UTC, Zenodo will be unavailable for 3-5 minutes to perform a storage cluster upgrade.

Published August 7, 2026

| Version v0.1

Preprint

Open

Valid Evidence Is Not Necessarily Comparable Evidence: Candidate-Conditioned Non-Invariance in LLM-Generated Code Verification

Authors/Creators

Kadri, Haitam<br>(Researcher)1

Show affiliations

1.

Independent Researcher

Description

Large language models are increasingly used to generate tests and other executable evidence for evaluating candidate code. This creates a comparative measurement problem when the evidence used to score a candidate is itself generated after observing that candidate.

This work studies candidate-conditioned comparative non-invariance using a controlled crossed design. Across 16 requirements spanning eight semantic families, changing which candidate was visible during test generation shifted the relative evaluation of the same fixed candidate pair and produced strict winner reversals in 5 of 16 requirements. The effect persisted under a valid-only analysis, while requirement-level analyses show that the aggregate direction should not be interpreted as universal across requirements.

A separate controlled defect-injection study shows the complementary benefit of candidate-aware verification: exposing the defective implementation increased targeted valid defect detection from 37.5% to 53.44%. These results distinguish diagnostic usefulness from comparative suitability.

The central conclusion is that specification validity is a per-item property and does not by itself confer cross-candidate comparability. Candidate-aware evidence can be useful for diagnosis, while shared evidence is preferable for direct candidate comparison.

The accompanying artifact contains frozen experimental configurations, execution results, integrity audits, analysis outputs, and reproducibility information.

Files

candidate-conditioned-noninvariance-v0.1-final.pdf

Files<br>(616.1 kB)

Name<br>Size

Download all

candidate-conditioned-noninvariance-v0.1-final.pdf

md5:8518ff2ac78e6cd06d62cb3c835e22fa

318.1 kB

Preview

Download

candidate-conditioned-noninvariance-v0.1-release.zip

md5:a1524815b6c78e976e238403f90b0f39

298.0 kB

Preview

Download

Views

Downloads

Show more details

All versions<br>This version

Views

Total views

Downloads

Total downloads

Data volume

Total data volume

0 Bytes<br>0 Bytes

More info on how stats are collected....

Versions

External resources

Indexed in

OpenAIRE

Communities

Keywords and subjects

Keywords

large language models

LLM evaluation

software engineering

code verification

test generation

coding agents

agent evaluation

AI evaluation

candidate selection

evaluation reliability

comparative evaluation

LLM-generated tests

comparative non-invariance

evidence comparability

Details

DOI

DOI Badge

DOI

10.5281/zenodo.21841560

Markdown

[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.21841560.svg)](https://doi.org/10.5281/zenodo.21841560)

reStructuredText

.. image:: https://zenodo.org/badge/DOI/10.5281/zenodo.21841560.svg<br>:target: https://doi.org/10.5281/zenodo.21841560

HTML

Image URL

https://zenodo.org/badge/DOI/10.5281/zenodo.21841560.svg

Target URL

https://doi.org/10.5281/zenodo.21841560

Resource type<br>Preprint

Publisher<br>Zenodo

Languages

English

Rights

License

Creative Commons Attribution 4.0 International

The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited.

Read more

Copyright

Copyright (C) 2026 Haitam Kadri.

Citation

Export

Technical metadata

Created

August 7, 2026

Modified

August 7, 2026

Jump up

This site uses cookies. Find out more on how we use cookies

Accept all cookies<br>Accept only essential cookies

candidate zenodo evidence conditioned evaluation https

Related Articles