When Verification Explores Too Far: Semantic Coverage and Validity in LLM-Generated Code Checks | Zenodo
Skip to main
You are using an outdated browser. Please upgrade your browser to improve your experience.
Published August 2, 2026
| Version v0.1
Publication
Open
When Verification Explores Too Far: Semantic Coverage and Validity in LLM-Generated Code Checks
Authors/Creators
Kadri, Haitam<br>(Researcher)1
Show affiliations
1.
Independent Researcher
Description
Large language models are increasingly used not only to generate code, but also to generate tests and other evidence intended to verify that code. This creates a methodological problem: a verifier may appear stronger when it explores behaviors beyond the public examples, while some of that apparent coverage may consist of requirements the verifier has invented rather than requirements supported by the specification. We study this problem in a controlled pilot in which an LLM generates executable verification challenges from a software requirement and public interface without observing the candidate implementation. We vary the generator's access to, and dependence on, public examples and separately measure challenge validity and semantic coverage. Challenges are labeled before condition identities are revealed, and validity is evaluated against a known-good implementation under a frozen requirement contract. Across two semantically distinct software bug families, we observe the same broad pattern: conditions that encourage greater exploration produce broader semantic coverage but lower challenge-level validity. In one replication case, the baseline condition achieved 75% validity and covered two useful semantic signatures, while the strongest exploration condition achieved 30% validity but covered nine useful signatures, including five not observed under the other conditions. Conversely, a structured condition achieved 100% validity while collapsing to a single semantic signature. A secondary validity-label audit on a frozen sample achieved 81.25% agreement (Cohen's kappa = 0.684) under evidence parity; annotator independence could not be verified retrospectively. These results are preliminary and do not establish a universal property of LLM verification. They instead identify a measurable evaluation problem: semantic novelty alone can overstate verification quality when the generated checks are not grounded in the requirement they are intended to verify.
Files
when_verification_explores_too_far_preprint_DOI.pdf
Files<br>(298.1 kB)
Name<br>Size
Download all
when_verification_explores_too_far_preprint_DOI.pdf
md5:877672cddff610c4a597ff5b44bc4136
298.1 kB
Preview
Download
Additional details
Related works
Is supplemented by
Software:
10.5281/zenodo.21755310
(DOI)
Views
Downloads
Show more details
All versions<br>This version
Views
Total views
Downloads
Total downloads
Data volume
Total data volume
0 Bytes<br>0 Bytes
More info on how stats are collected....
Versions
External resources
Indexed in
OpenAIRE
Communities
Keywords and subjects
Keywords
Large Language Models
Software Verification
Software Testing
Test Oracles
LLM
Evaluation
AI-Generated Code
Semantic Coverage
AI Agents
Details
DOI
DOI Badge
DOI
10.5281/zenodo.21758550
Markdown
[](https://doi.org/10.5281/zenodo.21758550)
reStructuredText
.. image:: https://zenodo.org/badge/DOI/10.5281/zenodo.21758550.svg<br>:target: https://doi.org/10.5281/zenodo.21758550
HTML
Image URL
https://zenodo.org/badge/DOI/10.5281/zenodo.21758550.svg
Target URL
https://doi.org/10.5281/zenodo.21758550
Resource type<br>Publication
Publisher<br>Zenodo
Languages
English
Rights
License
Creative Commons Attribution 4.0 International
The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited.
Read more
Copyright
Copyright (C) 2026 Haitam Kadri
Citation
Export
Technical metadata
Created
August 2, 2026
Modified
August 2, 2026
Jump up
This site uses cookies. Find out more on how we use cookies
Accept all cookies<br>Accept only essential cookies