AI 中文总结
本文批判性审视GEO可见性评分,提出框架明确提示语料库、权重等选择如何定义答案市场,并区分部分识别与规范性敏感性,强调引用不等于贡献。
AI 中文摘要
GEO(生成式引擎优化)可见性分数聚合了生成答案中的来源出现、引用或品牌提及。提示语料库选择了被评估的情境,而权重决定了它们的相对重要性。它们共同定义了一个“答案市场”,该市场不一定代表实际用户需求。提示措辞可以改变检索结果、竞争来源和生成的答案。评分随后需要识别感兴趣的出现、引用或提及。如果语言模型执行此任务,其指令可以改变对未变答案所赋予的分数。我们的批判性综述考察了这些选择如何帮助定义GEO分数所衡量的内容。它借鉴了关于指标是否衡量预期现象、总调查误差和信息检索评估的研究。该框架规定了情境标注、提示表述、执行条件、权重和评分规则。当权重未知或尚待选择时,框架报告可接受的分数集合。它区分了与数据和目标人群假设兼容的值(部分识别)与跨加权惯例的变化(规范性敏感性)。仅凭引用并不能确定来源的贡献。本文定义了一种在受控文献语境中,对有无来源生成的答案进行比较的方法,这不同于对具有竞争来源的完整引擎进行干预。该框架由可复现的计算支持。未报告新实验;其一般经验有效性仍有待评估。
英文摘要
GEO (generative engine optimization) visibility scores aggregate source appearances, citations, or brand mentions in generated answers. The prompt corpus selects the situations evaluated, while weights determine their relative importance. Together they define an "answer market" that need not represent actual user demand. Prompt wording can alter retrieval, competing sources, and generated answers. Scoring then requires identifying the appearances, citations, or mentions of interest. If a language model performs this task, its instruction can change the score assigned to an unchanged answer. Our critical survey examines how these choices help define what a GEO score measures. It draws on research into whether indicators measure the intended phenomenon, total survey error, and information retrieval evaluation. The framework specifies situation annotation, prompt formulations, execution conditions, weights, and scoring rules. When weights are unknown or remain to be chosen, the framework reports sets of admissible scores. It distinguishes values compatible with data and assumptions about a target population (partial identification) from variation across weighting conventions (normative sensitivity). A citation alone does not establish a source's contribution. The article defines a comparison of answers generated with and without a source in a controlled documentary context, distinct from an intervention on the full engine with competing sources. The framework is supported by reproducible calculations. No new experiments are reported; its general empirical validity remains to be assessed.
Comments16 pages, 9 tables; critical survey with 44 references and methodological analysis; ancillary reading register, aggregate data, and code for reproducing the calculations included