arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越二元记忆:面向多方口语对话的交互感知多模态记忆与自适应智能体检索

Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations

Wenxu Jia, Xize Cheng, Zihan Zhang, Dongjie Fu, Linjun Li, Wenshi Chen, Yangyang Wu, Tao Jin

arXiv 2609.32522首次发表:更新:

AI 中文总结

针对多方口语对话长期记忆缺失问题,提出交互感知多模态记忆框架VoxPolyMem,结合增量说话人识别与记忆层级,并引入EG-GRPO优化检索,在VoxPolyBench等基准上显著超越基线。

AI 中文摘要

长期记忆使智能体能够跨会话积累信息并进行推理,然而现有研究主要聚焦于二元文本或图像-文本对话,多方口语对话的长期记忆尚未得到充分探索。该场景要求保留对话内容、识别跨会话的参与者,并记录谁对谁说话。为此,我们提出VoxPolyMem,一个交互感知的多模态记忆框架,将增量式说话人识别与包含交互记忆、事实记忆和参与者档案的记忆层级相结合。我们将检索建模为序列决策过程,智能体基于累积证据重写查询并选择检索工具和记忆层,以解决信息缺口。我们进一步引入证据增益GRPO(EG-GRPO),利用逐轮信用分配鼓励互补性证据获取。我们还构建了VoxPolyBench,用于评估多方口语对话中的记忆演化、个性化回答、记忆检索与推理,以及交互推理与归因。VoxPolyMem在VoxPolyBench上取得85.0的总分,超过最强评估基线23.6分。在Mem-Gallery和H2HMem-Multi上,其得分分别为89.6和74.4,均超过最强公开记忆基线8分以上。这些结果凸显了其在多方多模态交互中实现持久、个性化辅助的潜力。代码和数据集可在以下网址获取:https URL

英文摘要

Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational content, identifying participants across sessions, and retaining who speaks to whom. To this end, we propose VoxPolyMem, an interaction-aware multimodal memory framework combining incremental speaker identification with a memory hierarchy comprising interaction memory, fact memory, and participant profiles. We formulate retrieval as sequential decision-making, where an agent rewrites queries and selects retrieval tools and memory layers based on accumulated evidence to address information gaps. We further introduce Evidence-Gain GRPO (EG-GRPO), which uses round-wise credit assignment to encourage complementary evidence acquisition. We also construct VoxPolyBench to evaluate memory evolution, personalized answering, memory retrieval and reasoning, and interaction reasoning and attribution in multi-party spoken conversations. VoxPolyMem achieves an overall score of 85.0 on VoxPolyBench, surpassing the strongest evaluated baseline by 23.6 points. On Mem-Gallery and H2HMem-Multi, it scores 89.6 and 74.4, respectively, exceeding the strongest evaluated public memory baselines by over 8 points each. These results highlight its potential for persistent, personalized assistance in multi-party multimodal interactions. Code and datasets are available at https://voxpolymem.github.io/VoxPolyBench/demo/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑