[2608.15896] When Search Eats the Web: A Model of Corpus Erosion under Generative Extraction
Skip to main content
Search arXiv
Press Enter to search · Advanced search
-->
Computer Science > Computer Science and Game Theory
arXiv:2608.15896 (cs)
[Submitted on 16 Aug 2026]
Title:When Search Eats the Web: A Model of Corpus Erosion under Generative Extraction
Authors:Sylvain Peyronnet<br>View a PDF of the paper titled When Search Eats the Web: A Model of Corpus Erosion under Generative Extraction, by Sylvain Peyronnet
View PDF<br>HTML (experimental)
Abstract:Generative search engines (GSEs) answer user queries directly from crawled web content. The capture of value from the corpus without a visit returned to the source (we call this capture extraction) diverts the traffic that finances content production. In response, publishers may restrict crawler access to their websites. In this paper, we model the crawlable corpus as a common-pool resource: the crawlable commons. It is described by three quantities: volume, average quality, and lifetime. Under two types of responses of publishers we prove that extraction degrades all three at once: publishers opt out, renewal loses its funding, and content becomes more perishable. After a given erosion threshold, the corpus goes extinct. A myopic GSE can cross this threshold, a long-run oriented GSE stays below it. We extend our model to several competing engines and prove, under a concavity condition on the steady-state value of the commons, that the symmetric equilibrium extraction rate is nondecreasing in their number and converges to the threshold. Adding users who strictly prefer direct answers, the assumption most favorable to extraction, we prove that the socially optimal extraction rate lies strictly below the erosion threshold, and no higher than the single engine's sustainable optimum. Finally, we discuss seven survival mechanisms.
Comments:<br>19 pages including a 1 page appendix
Subjects:
Computer Science and Game Theory (cs.GT); Information Retrieval (cs.IR)
MSC classes:<br>91A80 (Primary), 91B76 (Secondary)
ACM classes:<br>H.3.3; J.4
Cite as:<br>arXiv:2608.15896 [cs.GT]
(or<br>arXiv:2608.15896v1 [cs.GT] for this version)
https://doi.org/10.48550/arXiv.2608.15896
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Submission history<br>From: Sylvain Peyronnet [view email]<br>[v1]<br>Sun, 16 Aug 2026 19:03:13 UTC (41 KB)
Full-text links:<br>Access Paper:
View a PDF of the paper titled When Search Eats the Web: A Model of Corpus Erosion under Generative Extraction, by Sylvain Peyronnet<br>View PDF<br>HTML (experimental)<br>TeX Source
view license
Current browse context:
cs.GT
next >
new<br>recent<br>| 2026-08
Change to browse by:
cs<br>cs.IR
References & Citations
NASA ADS<br>Google Scholar
Semantic Scholar
export BibTeX citation<br>Loading...
BibTeX formatted citation
×
loading...
Data provided by:
Bookmark
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
TXYZ.AI (What is TXYZ.AI?)
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .
Which authors of this paper are endorsers? |<br>Disable MathJax (What is MathJax?)
Major funding support from