arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19059cs.CLcs.AI

MIRAGE:对话状态如何影响多模态个人代理中的历史证据使用

MIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents

Yu Liu, Wenxiao Zhang, Cheng Hu, Cong Cao, Fangfang Yuan, Xinyu Wang, Jin B. Hong, Yanbing Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出MIRAGE框架,通过控制对话状态变化评估多模态个人代理的历史证据使用,发现压缩前后存在非单调失败机制,强调需在状态变化下评估而非仅看结果正确性。

中文摘要 AI 辅助

多模态大语言模型(MLLM)代理越来越多地被用作长期任务的个人助理。其效用取决于连续性:代理必须跨对话、文件和工作区状态检索并使用早期证据。然而,即使访问历史记录的能力已退化,代理仍可能生成看似合理的答案,导致仅基于结果的评估高估了真实的证据使用情况。我们提出了MIRAGE(多模态交互检索、归因和接地评估),一项在对话状态变化下对多模态个人代理中历史证据使用进行受控研究。MIRAGE在保持证据对象、问题和评分不变的情况下,仅改变对话状态,并评估代理能否确定可回答性、恢复正确的来源并据此作答。在七个前沿和开放权重多模态骨干模型上,我们发现:1)压缩前深度和压缩后延续形式构成了不同的、非单调的失败机制,而非单一的退化曲线;2)开放权重模型严重依赖上下文连续性,当出处失败时不愿自发切换到工具介导的检索;3)对于符合工具使用的模型,检索压力在深度压缩前状态下改善了来源归因,但在压缩后持续退化,此时存储的证据已经退化。这些发现表明,历史证据使用应在状态变化下进行评估,而非仅从结果正确性推断。

英文摘要

Multimodal large language model (MLLM) agents are increasingly used as personal assistants for long-running tasks. Their utility depends on continuity: agents must retrieve and use earlier evidence across dialogue, files, and workspace state. However, agents can generate plausible answers even when access to that history has degraded, causing outcome-only evaluation to overestimate true evidence use. We present MIRAGE (Multimodal Interaction Retrieval, Attribution, and Grounding Evaluation), a controlled study of historical evidence use under conversation-state variation in multimodal personal agents. MIRAGE holds evidence objects, questions, and scoring fixed while varying only conversation state, and evaluates whether an agent can determine answerability, recover the correct source, and answer from it. Across seven frontier and open-weight multimodal backbones, we find that: 1) pre-compaction depth and post-compaction continuation form distinct, non-monotonic failure regimes rather than a single degradation curve; 2) open-weight models rely heavily on context continuity and are reluctant to spontaneously switch to tool-mediated retrieval when provenance fails; and 3) retrieval pressure improves source attribution in deep pre-compaction states for tool-compliant models, but consistently regresses after compaction, where stored evidence has already degraded. These findings show that historical evidence use should be evaluated under state variation, rather than inferred from outcome-only correctness.

发表机构

  • Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
  • School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
  • Department of Computer Science and Software Engineering, The University of Western Australia(西澳大学计算机科学与软件工程系)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑