arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越最终提示:衡量对话内上下文对AI回答的影响

Beyond the Final Prompt: Measuring the Effect of Within-Conversation Context on AI Answers

Benjamin Tannenbaum

arXiv 2608.02556首次发表:更新:

AI 中文总结

该研究通过180个多轮对话实验,发现AI回答受对话内上下文影响,完整对话生成的回答与仅用最终消息生成的回答在44.7%案例中有实质性差异,压缩前缀无法完全替代完整上下文。

AI 中文摘要

在AI系统评估中,通常将孤立的最终用户消息视为查询。然而在对话中,可执行请求可能分布在前面的轮次中。我们直接测试被省略的对话内上下文是否会改变回答。我们从受监管的商业语料库和公开的PRISM数据集中采样了180个英文多轮对话,对于每个对话,我们保持最终用户消息和所请求的回答模型不变,同时生成三个回答:一个来自完整的带角色标签的对话,一个仅来自最终消息,还有一个来自最终消息加上前缀仅重建(上限为160词)。一个单独请求的评判模型在随机标签下评估回答。预先指定的主要终点是可能改变用户行为的实质性差异,而非风格或细节上的差异。对符合条件的队列进行逆概率加权后,完整对话与孤立最终消息的回答在44.7%的案例中存在实质性差异(95%自助法置信区间为33.8%至56.1%)。完整对话的回答在0至4分的请求满意度量表上得分高出0.49分(0.32至0.67)。添加压缩前缀将实质性差异率降至30.8%(20.2%至42.1%),减少了13.9个百分点(4.9%至24.1%),并将平均满意度差距降至0.01分(-0.12至0.13)。但压缩并不等同于完整对话上下文:近三分之一的回答仍存在实质性差异。对48个案例进行的顺序交换重复实验显示,主要决策的一致性为91.7%,kappa值为0.83。该研究关注同一对话中的前面轮次,未测试跨不同对话的持久记忆。

英文摘要

An isolated final user message is often treated as the query in evaluations of AI systems. In a conversation, however, the actionable request may be distributed across preceding turns. We directly test whether that omitted within-conversation context changes answers. For each of 180 English multi-turn conversations sampled from a governed commercial corpus and the public PRISM dataset, we hold the final user message and requested answer model constant while generating three answers: one from the full role-labelled conversation, one from the final message alone, and one from the final message plus a prefix-only reconstruction capped at 160 words. A separately requested judge model evaluates answers under randomized labels. The prespecified primary endpoint is a material difference that could change what the user does, rather than a difference in style or detail. After inverse-probability weighting to the eligible cohorts, the full-conversation and isolated-final answers differ materially in 44.7% of cases (95% bootstrap CI 33.8% to 56.1%). Full-conversation answers score 0.49 points higher on a 0 to 4 request-satisfaction scale (0.32 to 0.67). Adding the compressed prefix reduces the material-difference rate to 30.8% (20.2% to 42.1%), a 13.9-point reduction (4.9% to 24.1%), and reduces the mean satisfaction gap to 0.01 points (-0.12 to 0.13). Yet compression is not equivalent to the complete dialogue context: almost one third of answers remain materially different. An order-swapped repeat on 48 cases yields 91.7% agreement and kappa = 0.83 for the primary decision. The study concerns preceding turns in the same conversation and does not test persistent memory across separate conversations.

Comments8 pages, 3 figures, 2 tables. Companion to arXiv:2607.22392

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑