arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07725cs.AI

PERSIST:全双工口语对话中跨会话的谁-什么-何时记忆

PERSIST: Who-What-When Memory Across Sessions for Full-Duplex Spoken Dialogue

Achira Lin, Siyuan Hou, Wenyi Yu, Xinnian Zhao, Haoyu Niu, Wang Geng, Longshuai Xiao, Shihai Xiao, Mangsuo Zhao, Chao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对多用户共享语音助手需跨会话记忆的问题,提出PERSIST系统,显式建模谁-什么-何时,用3W联合评分检索,并复用中间表示将延迟降至7.03毫秒,在SpokenTrace基准上端到端准确率达85.08%。

中文摘要 AI 辅助

现代语音助手可能被多个用户共享,并且应当能够回答关于早期对话的问题,例如“我原本计划什么时候离开?”,或基于过去的交互调整其行为以适应个体用户。这需要的不只是检索主题相似的段落:助手必须识别当前说话者,恢复相关的过去状态,并将其与后来的修订区分开来。我们提出PERSIST,一个用于多会话、多说话者口语对话的持久记忆系统,它显式地建模谁、什么和何时。PERSIST将跨会话历史结构化为可读的事件记录,并通过一种3W联合评分机制进行检索,该机制结合了语义内容、声学说话者身份和时间状态。对于实时全双工交互,PERSIST进一步重用对话主干的中间表示,避免查询音频重新编码,并将检索延迟从578.42毫秒降低到7.03毫秒。我们还引入了SpokenTrace,一个诊断性基准,它沿记忆任务和说话者查询类型对评估进行分解,揭示了在回忆、说话者归因和时间状态跟踪方面的失败。在SpokenTrace上,PERSIST实现了85.08%的端到端任务准确率,并将全支持EM@3从使用BGE-large的49.01%提高到82.10%。

英文摘要

Modern voice assistants may be shared by multiple users and should be able to answer questions about earlier conversations such as "When did I originally plan to leave?" or adapt their behavior to individual users based on past interactions. This requires more than retrieving a topically similar passage: the assistant must identify the current speaker, recover the relevant past state, and distinguish it from later revisions. We present PERSIST, a persistent memory system for multi-session, multi-speaker spoken dialogue that explicitly models Who, What, and When. PERSIST structures cross-session histories into readable event records and retrieves them with a 3W joint scoring mechanism that combines semantic content, acoustic speaker identity, and temporal state. For real-time full-duplex interaction, PERSIST further reuses intermediate representations from the dialogue backbone, avoiding query-audio re-encoding and reducing retrieval latency from 578.42 ms to 7.03 ms. We also introduce SpokenTrace, a diagnostic benchmark that factorizes evaluation along memory tasks and speaker-query types, exposing failures in recall, speaker attribution, and temporal-state tracking. On SpokenTrace, PERSIST achieves 85.08% end-to-end task accuracy and improves all-support EM@3 from 49.01% with BGE-large to 82.10%.

发表机构

  • Tsinghua University(清华大学)
  • Huawei Technologies Ltd.(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑