arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

什么驱动生产级大语言模型中的引用?对一万个网页上两百万条AI引用的观察性多方法研究

What Drives Citations in Production Large Language Models? An Observational Multi-Method Study of Two Million AI Citations Across Ten Thousand Web Pages

Ben Moore, Liam Dunne

arXiv 2609.35077首次发表:更新:

发表机构

Discovered Labs(Discovered Labs)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究基于两百万条商业LLM引用和一万个网页,发现提示-内容对齐是引用频率的主导预测因子,AEO清单效应在控制领域后反转,领域权威影响显著,并发布分析流程。

AI 中文摘要

生产级大语言模型在生成答案的同时会检索并引用网页,然而预测引用频率的页面级特征仍未得到充分刻画。我们开展了一项观察性研究,涵盖六个月期间来自四个商业引擎(ChatGPT、Claude、Google AI、Gemini)的约两百万条LLM引用,并将其与来自十九个B2B SaaS工作空间的一万个爬取页面进行关联。我们使用一个九方法共识框架测试了六十多个特征,该框架结合了混合效应回归与领域固定效应、FDR校正、稳定性选择Lasso、双重机器学习、广义加性模型以及时间保持复制。四项发现通过了所有检验。第一,提示-内容对齐(页面词元与完整工作空间提示语料库(包括非引用提示)之间的Jaccard重叠)是主导性的页面级预测因子(beta = +0.37,95%置信区间[+0.33,+0.41],q ~ 10^-73)。第二,标准的AEO检查清单(FAQ块、结构化数据、Core Web Vitals)在合并数据中显示出正向效应,但在应用领域固定效应后这些效应反转或缩减为零:这是辛普森悖论,对AEO文献具有实际后果。第三,领域级AI权威度的平均绝对SHAP值比最强的非对齐页面级特征高出六倍。我们发布了分析流程作为方法论贡献。

英文摘要

Production large language models retrieve and cite web pages alongside generated answers, yet the page-level features that predict citation frequency remain poorly characterised. We present an observational study of approximately 2 million LLM citations from four commercial engines (ChatGPT, Claude, Google AI, Gemini) over six months, joined to 10,000 crawled pages from nineteen B2B SaaS workspaces. Sixty-plus features are tested using a nine-method consensus framework combining mixed-effects regression with domain fixed effects, FDR correction, stability-selection Lasso, double machine learning, generalised additive models, and temporal hold-out replication. Four findings survive all checks. First, prompt-content alignment (Jaccard overlap between page tokens and the full workspace prompt corpus, including non-citing prompts) is the dominant page-level predictor (beta = +0.37, 95% CI [+0.33, +0.41], q ~ 10^-73). Second, the standard AEO checklist (FAQ blocks, structured data, Core Web Vitals) shows positive effects in pooled data that reverse or collapse to zero once domain fixed effects are applied: Simpson's paradox with practical consequences for the AEO literature. Third, domain-level AI authority exceeds the strongest non-alignment page-level feature by a factor of six in mean absolute SHAP value. We release the analytic pipeline as a methodological contribution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑