发表机构
Harbin Institute of Technology, Shenzhen; AgiBot; Tsinghua Shenzhen International Graduate School, Tsinghua University(哈尔滨工业大学(深圳); AgiBot; 清华大学深圳国际研究生院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多模态交互中历史使用评估的空白,构建了含7000个任务的ReTurn基准,评估13类模型的选择性历史使用能力,发现其准确率下降且监督适应增益有限。
AI 中文摘要
可靠的多模态交互依赖于对话历史的选择性使用:早期问题可能仍具相关性,但其先前的回答已过时;而当前请求可能需要依赖历史证据,尽管存在冲突的新观测结果。现有多轮评估很少将这些历史使用需求与潜在的问题难度分开。为解决这一差距,我们引入ReTurn,这是一个包含7000个基础任务的基准,涵盖视觉和音频证据,用于评估选择性历史使用。对于承载任务的历史,Reconfirm/Reground要求将历史问题应用于当前媒体,同时改变历史一致性;对于承载证据的历史,Retrieve/Rebind要求使用历史媒体回答当前问题,同时改变当前媒体的竞争性。每对任务都保留目标问题、媒体和答案。任务支持开放式和多项选择评估,匹配的单轮对应任务作为可答性参考。在13个全模态、视觉-语言和音频-语言模型中,模型级别的开放式准确率中位数从直接输入时的93.7%降至对话中的72.3%。行为探测显示,高问题召回率可与较弱的任务应用共存,而竞争性媒体可将答案从历史目标中转移。监督适应仅产生部分增益。ReTurn提供了一个受控框架,用于评估多模态模型是否选择并使用每个请求所需的历史信息。
英文摘要
Reliable multimodal interaction depends on selective use of conversational history: an earlier question may remain relevant while its previous answer is outdated, whereas a current request may depend on historical evidence despite conflicting new observations. Existing multi-turn evaluations rarely separate these history-use demands from underlying question difficulty. To address this gap, we introduce ReTurn, a benchmark of 7,000 base tasks spanning visual and audio evidence for evaluating selective history use. For task-carrying history, Reconfirm/Reground require applying a historical question to current media while varying historical agreement; for evidence-carrying history, Retrieve/Rebind require answering a current question using historical media while varying current-media competition. Each pair preserves the target question, media, and answer. Tasks support open-ended and multiple-choice evaluation, with matched single-turn counterparts serving as answerability references. Across 13 omni-modal, vision-language, and audio-language models, median model-level open-ended accuracy falls from 93.7% with direct input to 72.3% in conversation. Behavioral probes show that high question recall can coexist with weaker task application, while competing media can redirect answers away from historical targets. Supervised adaptation yields only partial gains. ReTurn provides a controlled framework for assessing whether multimodal models select and use the historical information required by each request.