arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00570cs.AIcs.LG

VoiceLongMemEval:助手能否记住你的声音?

VoiceLongMemEval: Do Assistants Remember How You Sounded?

Ramit Pahwa, Parivesh Priye, Apoorva Beedu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对现有对话基准忽略人机交互副语言信息的问题,提出VoiceLongMemEval基准,发现副语言元数据可提升模型准确率,原生音频模型能直接从语音提取相关线索。

中文摘要 AI 辅助

随着多智能体架构和大型语言模型规模不断扩大,已部署的AI助手越来越需要对长期、连续、多轮会话历史进行推理。当前基准将对话历史评估为长距离信息检索、时间推理或知识更新,却忽略了人机交互的基本动态——即说话方式。为填补这一空白,我们提出VoiceLongMemEval(VLME)基准,其中每个答案都依赖于会话轮次附带的副语言元数据(情绪标签、韵律描述符和声音事件),仅从文本中无法恢复这些信息。每个条目都经过三阶段对抗性筛选,确保强大的语言模型仅在给定文本转录时会失败。对前沿开源权重模型的评估显示存在普遍的情感差距;提供文本轨道的副语言元数据可使准确率提升0.09至0.38(当提示带有证据提示时,准确率为0.61至0.69),而标准自动语音识别(ASR)流程会系统性丢弃该信号。此外,原生音频模型可直接从语音中成功提取这些线索(准确率为0.354至0.412,对比盲态模型的0.325)。代码和数据集将在论文接收后提供。

英文摘要

With the growing scale of multi-agent architectures and large language models, deployed AI assistants are increasingly tasked with reasoning over long, continuous, multi-session conversation histories. Current benchmarks evaluate this dialogue history as information retrieval over long horizon, temporal reasoning, or knowledge updates, while crucially ignoring the fundamental dynamics of human-agent interaction, i.e. how they said it. To address this gap, we present VoiceLongMemEval (VLME) benchmark, where every answer depends on paralinguistic metadata (emotion labels, prosody descriptors, and voice events) attached to conversational turns, which is otherwise unrecoverable from the words alone. Every item passes a three-stage adversarial gate, ensuring that a strong language model fails when given only the transcript. Evaluating leading frontier and open-weight models reveals a pervasive affect gap; providing text-track paralinguistic metadata yields a 0.09 to 0.38 accuracy boost (0.61 to 0.69 when prompted with evidence hints), while standard ASR pipelines systematically discard this signal. Additionally, audio-native models successfully extract these cues directly from speech (0.354 to 0.412 vs. 0.325 blind). Code and dataset will be made available upon acceptance.

↑