用引擎评分:检索曝光、跨引擎差异与引擎无关GEO评分的局限性
Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores
查看机构详情
- Aiso Boost Ltd.(Aiso Boost 有限公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究通过审计四个生成式引擎,发现跨引擎引文重叠极低,表明引擎无关评分只能衡量页面质量,无法替代端到端可见性,需分别报告曝光、选择等不同量。
中文摘要 AI 辅助
近期研究探讨是否可以用确定性的、与引擎无关的页面评分来近似生成式引擎的可见性。我们将此类评分可能混淆的两个阶段区分开来:在实时引擎上的曝光,以及基于曝光的引文选择。在对ChatGPT、Microsoft Copilot、Google和Perplexity的观察性审计中,15个固定商业提示在2026年6月6日产生了589条引文观察,对应528个唯一URL和356个域名。相同提示下跨引擎的URL重叠极小:平均成对Jaccard相似度为0.0079,中位数为零,84.9%的引擎对没有共享任何被引URL。在四个引擎上均观察到的十个提示中,平均精确URL Jaccard为0.0072。一个匹配大小的超几何基线(保留每个提示的四引擎URL全集和每个引擎的列表长度)预测值为0.1272,因此观察到的重叠仅为该基线的5.7%;零URL重叠发生在86.7%的比较中,而预期为12.3%。前五精确URL重叠在所有60个成对比较中均为零。单个引擎仅捕获了四引擎URL并集的11.4%-42.6%,且96.4%的观察URL仅出现在一个引擎中。另一次6月5日至6日的同引擎比较发现平均URL集周转率为67.0%。这些结果并未否定与引擎无关的页面评分;它们明确了其估计目标。不使用实时引擎计算的评分可以估计页面质量或查询-页面匹配度,而端到端可见性还额外依赖于引擎特定的曝光和选择。因此,我们主张将页面匹配度、观察到的曝光、条件选择和最终可见性作为不同的量分别报告。
英文摘要
Recent work asks whether generative-engine visibility can be approximated with deterministic, engine-free page scores. We separate two stages such scores can conflate: exposure to a live engine and citation selection conditional on exposure. In an observational audit of ChatGPT, Microsoft Copilot, Google, and Perplexity, 15 fixed commercial prompts produced 589 citation observations on 6 June 2026, corresponding to 528 unique URLs and 356 domains. Same-prompt cross-engine URL overlap was extremely small: mean pairwise Jaccard similarity was 0.0079, the median was zero, and 84.9% of engine pairs shared no cited URL. On the ten prompts observed on all four engines, mean exact-URL Jaccard was 0.0072. A matched-size hypergeometric baseline preserving each prompt's four-engine URL universe and each engine's list length predicts 0.1272, so observed overlap was only 5.7% of that baseline; zero URL overlap occurred in 86.7% of comparisons versus 12.3% expected. Top-five exact-URL overlap was zero in all 60 pairwise comparisons. A single engine captured only 11.4%-42.6% of the four-engine URL union, and 96.4% of observed URLs appeared in only one engine. A separate 5-to-6 June same-engine comparison found 67.0% mean URL-set turnover. These results do not invalidate engine-free page scoring; they identify its estimand. A score computed without a live engine can estimate page quality or query-page fit, while end-to-end visibility additionally depends on engine-specific exposure and selection. We therefore argue for reporting page fit, observed exposure, conditional selection, and final visibility as distinct quantities.