arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当搜索吞噬网络:生成式提取下的语料库侵蚀模型

When Search Eats the Web: A Model of Corpus Erosion under Generative Extraction

Sylvain Peyronnet

arXiv 2608.15896首次发表:更新:

AI 中文总结

本文将可爬取语料库建模为公共池资源,分析生成式搜索引擎的提取行为对语料库规模、质量和生命周期的侵蚀效应,证明其与发布者应对方式的关系,并提出生存机制。

AI 中文摘要

生成式搜索引擎(GSE)会直接从爬取的网络内容中响应用户查询。在不返回对来源网站的访问的情况下从语料库中获取价值(我们将此过程称为提取),会分流为内容生产提供资金的流量。作为应对,发布者可能会限制爬虫对其网站的访问。在本文中,我们将可爬取的语料库建模为公共池资源:可爬取公地,它由三个量描述:规模、平均质量和生命周期。在发布者的两种应对类型下,我们证明提取会同时降低这三个量:发布者选择退出、内容更新失去资金支持,且内容变得更易失效。在达到给定的侵蚀阈值后,语料库会灭绝。短视的GSE会越过该阈值,而着眼长期的GSE则会保持在阈值以下。我们将模型扩展到多个相互竞争的搜索引擎,并在公地稳态价值的凹性条件下证明,对称均衡提取率随搜索引擎数量增加而不降低,并收敛至该阈值。加入严格偏好直接答案的用户(这是对提取最有利的假设)后,我们证明社会最优提取率严格低于侵蚀阈值,且不超过单个搜索引擎的可持续最优值。最后,我们讨论了七种生存机制。

英文摘要

Generative search engines (GSEs) answer user queries directly from crawled web content. The capture of value from the corpus without a visit returned to the source (we call this capture extraction) diverts the traffic that finances content production. In response, publishers may restrict crawler access to their websites. In this paper, we model the crawlable corpus as a common-pool resource: the crawlable commons. It is described by three quantities: volume, average quality, and lifetime. Under two types of responses of publishers we prove that extraction degrades all three at once: publishers opt out, renewal loses its funding, and content becomes more perishable. After a given erosion threshold, the corpus goes extinct. A myopic GSE can cross this threshold, a long-run oriented GSE stays below it. We extend our model to several competing engines and prove, under a concavity condition on the steady-state value of the commons, that the symmetric equilibrium extraction rate is nondecreasing in their number and converges to the threshold. Adding users who strictly prefer direct answers, the assumption most favorable to extraction, we prove that the socially optimal extraction rate lies strictly below the erosion threshold, and no higher than the single engine's sustainable optimum. Finally, we discuss seven survival mechanisms.

Comments19 pages including a 1 page appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑