arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VoiceMem:用于实时交互的双脑流式记忆系统

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

Zhifei Xie, Jiaqi Lang, Ze An, Yifan Zhao, Dongchao Yang, Kai Li, Ziyang Ma, Mingbao Lin, Chunyan Miao, Shuicheng Yan

arXiv 2608.26005首次发表:更新:

发表机构

Nanyang Technological University; National University of Singapore; Tsinghua University; The Chinese University of Hong Kong(南洋理工大学; 新加坡国立大学; 清华大学; 香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出VoiceMem双脑流式记忆架构,构建完整流水线实现记忆感知SLM的训练、评估与部署,其在准确性、情感个性化及实时性上优势显著,为实时语音交互提供实用记忆基础。

AI 中文摘要

对话系统(如双工语音语言模型SLMs)仍缺乏流式、准确且共情的记忆系统作为核心。本文提出VoiceMem,一种简单的记忆架构,包含并行的信息左脑、情感右脑及流式记忆I/O机制。我们进一步构建了完整的流水线,用于记忆感知SLM训练、长视野评估及可替换记忆后端的解耦部署。实验与实际部署显示三大优势:i)准确性:在Top-5检索下,左脑性能超越Mem0等经典系统,Top-200时领先近30个百分点;ii)情感与个性化:右脑具备短/长视野情感归因及双节点角色建模,在三个角色基准上达SOTA性能,综合得分较此前最优系统提升4.29个百分点;iii)实时与低成本:VoiceMem完成检索耗时134ms,远在标准VAD延迟范围内,既无额外对话延迟,又保持高准确率与低成本。这些结果表明,VoiceMem为实时、个性化且具情感感知的语音交互提供了实用的记忆基础。

英文摘要

Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.

Comments18 pages, 9 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑