发表机构
University of Melbourne; University of New South Wales(墨尔本大学; 新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VoxMem提出一个多模态记忆基准,通过声学证据与记忆操作分类法,在34,743个口语会话上评估15个大型音频语言模型,发现模型在保留说话内容上远优于其他声学信息,且复杂操作下性能显著下降。
AI 中文摘要
口语对话系统必须从先前的交互中恢复信息(即记忆),然而语音中的相关信息不仅包括所说的内容,还包括说话者是谁、说话的方式以及可听到的环境声音,这些信息仅存在于音频信号中,无法从文本转录中恢复。除了记忆什么之外,记忆还要求多样化的操作:检索单一事实、跨轮次整合证据、跟踪不断变化的状态。真实交互进一步跨会话展开,这意味着信息在多个不同片段中累积,而非单一连续录音。现有基准在三个维度上均存在不足:它们主要关注词汇内容,采用有限且临时的记忆操作,并将记忆视为单会话问题。我们认为,原则性的记忆评估需要联合刻画要保留的声学证据以及对其应用的操作,并引入一个沿这两个轴线的分类法。基于此分类法,我们提出了VoxMem:包含3,196个评估实例,跨越34,743个口语会话(177小时),涵盖四种声学证据类型(语音语义、说话者身份、副语言线索、环境声音)与四种记忆操作(信息提取、跨会话推理、时间跟踪和答案拒绝),基于多会话历史构建,并在8K至64K token的上下文预算中分层。评估15个LALM,在32K时没有模型超过40%。模型对所说的内容保留远好于对说话者身份、说话方式或可听环境声音的保留,这一差距在复杂操作中扩大,随历史长度增长,并在不同证据类型中表现为定性不同的失败模式。VoxMem旨在为衡量和推动口语对话记忆全范围进展提供基础。
英文摘要
Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be recovered from a transcript. Beyond what to remember, memory also demands diverse operations: retrieving a single fact, integrating evidence across turns, tracking an evolving state. Real interactions further unfold across sessions, meaning information accumulates across distinct episodes rather than a single continuous recording. Existing benchmarks fall short on all three dimensions: they focus primarily on lexical content, adopt limited and ad hoc memory operations, and treat memory as a single-session problem. We argue that principled memory evaluation requires jointly characterizing the acoustic evidence to be retained and the operations applied to it, and introduce a taxonomy along these two axes. Building on this taxonomy, we present VoxMem: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal tracking, and answer refusal), grounded in multi-session histories and stratified across context budgets from 8K to 64K tokens. Evaluating 15 LALMs, no model exceeds 40% at 32K. Models retain what was said far better than who said it, how, or what was audible, a gap that widens for complex operations, grows with history length, and manifests as qualitatively distinct failure modes across evidence types. VoxMem aims to provide a foundation to measure and drive progress on the full scope of spoken conversational memory.
Commentsworking in process