AI 中文总结
研究多轮人工智能对话中请求状态演变,用对话条件请求状态取代潜在“意图”,分析其在用户轮次中的分布,通过复用规则和对话群组进行研究,揭示最终提示与历史请求状态差异,支持会话级测量,未估计历史对模型答案的因果效应。
AI 中文摘要
人工智能搜索评估通常将提示视为一个稳定的查询,可以单独进行计数、分类和重放。而对话使这种分析单元受到质疑,因为每个用户轮次都可以添加约束、修改假设、请求证据或引用先前建立的替代方案。我们用一个可观察的结构——对话条件请求状态来取代潜在的“意图”,并衡量该状态在用户轮次中的分布情况。分析复用了之前对话研究中的固定规则和受管理的群组:670个英语商业多轮对话采用发现-复制设计,以及来自1389名参与者的7463个公共PRISM对话。在商业语料库中,最终提示包含会话中独特用户端内容词汇的中位数为35.6%;在PRISM中,中位数为36.4%。最终提示在68.4%和74.3%的对话中分别包含至多一半的该词汇。更重要的是,透明规则在50.3%的商业对话和44.8%的PRISM对话中检测到历史中的至少一个请求状态维度,但在最终提示中未检测到。在有维度的对话中,最终提示仅在26.1%和26.2%的情况下重现完整观察到的维度集。同时,最终提示在17.9%和19.3%的情况下添加了一个先前未见过的维度,表明端点既不是总结也不仅仅是参考:它通常是另一个状态更新。长度匹配的空值表明低词汇覆盖率在很大程度上是轮次长度的结果,因此词汇结果被解释为信息可用性,而非语义漂移。分类结果支持人工智能搜索的会话级测量。它们没有估计历史对模型答案的因果效应。
英文摘要
AI-search evaluation commonly treats a prompt as a stable query that can be counted, classified, and replayed in isolation. A conversation makes that unit of analysis questionable: each user turn can add a constraint, revise an assumption, request evidence, or refer to alternatives established earlier. We replace latent "intent" with an observable construct, conversation-conditioned request state, and measure how that state is distributed across user turns. The analysis reuses frozen rules and the governed cohort of a preceding conversation study: 670 English commercial multi-turn conversations in a discovery-replication design and 7,463 public PRISM conversations from 1,389 participants. In the commercial corpus, the final prompt contains a median 35.6% of the session's unique user-side content vocabulary; in PRISM, the median is 36.4%. The final prompt contains at most half of that vocabulary in 68.4% and 74.3% of conversations, respectively. More importantly, transparent rules detect at least one request-state dimension in history but not in the final prompt in 50.3% of commercial conversations and 44.8% of PRISM conversations. Among dimension-bearing conversations, the final prompt reproduces the full observed dimension set in only 26.1% and 26.2%. At the same time, the final prompt adds a previously unseen dimension in 17.9% and 19.3%, showing that the endpoint is neither a summary nor merely a reference: it is often another state update. Length-matched nulls show that low lexical coverage is largely a consequence of turn length, so vocabulary results are interpreted as information availability, not semantic drift. The categorical results support session-level measurement for AI search. They do not estimate the causal effect of history on model answers.