发表机构
Estonian Entrepreneurship University of Applied Sciences (EUAS)(爱沙尼亚创业应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对6个大语言模型,通过300个问题-引擎单元实验发现,未检索的5个引擎品牌推荐未饱和,检索引擎已饱和,而引用领域积累持续上升,揭示了重复查询下品牌推荐与信息来源的不同特性。
AI 中文摘要
重复相同的购买问题是否会耗尽语言模型的品牌推荐取决于检索机制。在300个问题-引擎单元(50个问题、6个引擎、每个引擎15次运行、对1470个经裁决的组织进行开放提取)中,5个未进行网络搜索的引擎在第15次运行时,仍有86%-92%的单元会添加从未出现过的品牌,其中位推荐集包含15-31个组织;1个启用检索的引擎则已停止新增品牌(中位推荐集包含8个组织,仍有64%的单元在添加),这与之前4个深度单元的情况一致,这些单元的网络搜索运行在第10次时就已饱和。引用领域的积累在所有测试的时间范围内持续上升:4个深度单元在第24次运行时仍在添加领域,仅观察到Chao2下界估计值的59%-84%;检索引擎的广度单元中有44%在第15次运行时仍在添加领域。单次运行可展示5次运行品牌集的62%-77%,各引擎的中位问题会引出38个组织,其中中位15个组织仅在一个引擎中出现。所用估计量为精确稀疏法和Chao2丰富度;并行固定名单提取在相同响应上呈现平坦曲线,因此名单受限的追踪会产生平台期,而开放提取则消除了该现象。
英文摘要
Repeated-query audits must distinguish recovery of a collected set from completeness of possible outputs. We apply sample-based rarefaction to 4,500 responses from 50 buying questions, six configurations and 15 calls per cell. Historical-dictionary median ten-call recovery of the observed 15-call set ranges from 92.6% to 95.2%; re-adjudicating all 45,683 candidate strings changes this range to 89.5%-94.7%. Two blinded Gemini 3.1 Pro annotation roles assessed 600 complete answers, yielding micro F1 of 0.908 for canonical-name agreement and 0.975 for span-overlap agreement. This is AI-based evidence, without a human reference study. A separate matched roster analysis of 3,750 records per wave gives median single-call recovery of the observed five-call set of 80.0%-92.5% in February and 90.0%-100.0% in September, with question-subset dependence. Source accumulation also changes when API-returned hosts are restricted to those referenced by answer citation markers. These findings show that recovery percentages depend on extraction, question selection and the finite reference collection. They support explicit measurement definitions and sensitivity analyses, without establishing exhaustive repertoires, causal retrieval effects or a universal stopping rule.
Comments14 pages, 4 figures. Substantially revised finite-sample audit of repeated LLM queries with extraction sensitivity, a matched five-call comparison, and blinded AI-reference agreement analysis