发表机构
The University of Tokyo; UniConvo Inc.; University of Tsukuba; Hiroshima University; National Institute of Informatics(东京大学; UniConvo公司; 筑波大学; 广岛大学; 国立情报学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出对话捕获概念和轨迹级框架,证明单轮评估低估GEO效果,多轮中反馈项超线性增长,排名弱一致。
AI 中文摘要
生成式引擎优化(GEO)通过塑造内容以提高其在基于检索增强的大语言模型构建的答案引擎中被引用的可能性。GEO通常被评估为单轮属性:对于固定查询,评估器在单个答案中衡量来源的可见性。我们认为单个答案作为分析单元是不充分的。人机信息寻求形成一个闭环:代理的答案改变用户的信念,从而影响下一个问题,进而决定代理检索的内容。我们引入对话捕获现象,即早期被引用的来源在后续被再次引用的可能性显著增加。捕获通过机器侧通道(基于历史的检索)和人类侧通道(针对被捕获来源的后续问题)运作。我们将交互形式化为两层闭环系统,并推导出轨迹级概念:累积对话可见性;轨迹增益的直接/反馈分解;反馈项在机器侧和人类侧通道上的嵌套分割;捕获系数;复合比率;以及误排名诊断。利用强化过程(波利亚瓮)理论,我们证明在单轮评估下反馈项为零,且随着对话长度增加,GEO的累积收益呈超线性增长,同时捕获现象发展。一个模型推导的示例表明,反馈项可能超过直接项,复合比率在十轮内超过2,且单轮与轨迹排名仅弱一致(Kendall's τ = 0.4)。我们将人类通道与信息觅食、信任校准和贝叶斯说服联系起来,并讨论对答案引擎的设计启示。
英文摘要
Generative Engine Optimization (GEO) shapes content to increase its likelihood of being cited by answer engines built on retrieval-augmented large language models. GEO is typically evaluated as a single-turn property: for a fixed query, an evaluator measures a source's visibility in one answer. We argue that the single answer is an inadequate unit of analysis. Human-agent information seeking forms a closed loop: the agent's answer changes the user's beliefs and therefore the next question, which in turn determines what the agent retrieves. We introduce conversational capture, a phenomenon in which a source cited early becomes substantially more likely to be cited again. Capture operates through a machine-side channel, history-conditioned retrieval, and a human-side channel, follow-up questions directed toward the captured source. We formalize the interaction as a two-layer closed-loop system and derive trajectory-level constructs: cumulative conversational visibility; a direct/feedback decomposition of trajectory gain; a nested split of the feedback term into machine-side and human-side channels; a capture coefficient; a compounding ratio; and a misranking diagnostic. Using reinforcement-process (Pólya-urn) theory, we prove that the feedback term is zero under single-turn evaluation and that GEO's cumulative payoff grows superlinearly with conversation length while capture develops. A model-derived illustration shows that the feedback term can exceed the direct term, the compounding ratio exceeds two within ten turns, and single-turn and trajectory rankings agree only weakly (Kendall's $τ= 0.4$). We connect the human channel to information foraging, trust calibration, and Bayesian persuasion, and discuss design implications for answer engines.
Comments9 pages, 2 figures, 2 tables. In Proceedings of the 14th International Conference on Human-Agent Interaction (HAI '26), November 16-19, 2026, Osaka, Japan