arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22452cs.CLcs.HC

Vox-Infinity:基准测试长上下文语音语言模型的极限

Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models

  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Xize Cheng, Wenxu Jia, Chenyuhao Wen, Dongjie Fu, Zehan Wang, Xinyu Zhang, Tao Jin

AI总结:

Vox-Infinity是首个评估语音语言模型长上下文理解的基准,通过扩展轮次和时长揭示模型存在近因效应,即难以利用对话早期证据,为长上下文语音AI提供评测标准。

AI中文摘要:

长上下文理解仍然是大型语言模型面临的一个基本挑战,因为过长的输入往往会导致模型遗忘重要信息。这个问题在语音领域尤为突出,因为音频作为一种低压缩率模态,需要比文本多得多的嵌入来保留语义内容和声学线索。为了解决这一挑战,我们引入了\ extbf{Vox-Infinity},这是第一个专门设计用于评估语音语言模型长上下文理解的基准。Vox-Infinity沿着两个维度系统地扩展音频历史:轮次数量和轮次持续时间。它涵盖了具有不同交互结构和语义复杂性的多种代表性场景。至关重要的是,Vox-Infinity提供了明确的答案来源注释,并根据解决每个查询所需的历史上下文量对样本进行组织,从而实现精确且长度感知的评估。对七个代表性语音语言模型的广泛评估揭示了一个明显的整体近因效应:当支持答案的证据更接近查询时,模型通常能达到更高的准确率,但在检索和使用对话历史中更早出现的证据时则表现困难。案例和数据集可在以下网址获取:https URL。

英文摘要:

Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient information. This issue is even more pronounced in the speech domain, where audio, as a low-compression modality, requires substantially more embeddings than text to preserve both semantic content and acoustic cues. To address this challenge, we introduce \textbf{Vox-Infinity}, the first benchmark specifically designed to evaluate long-context understanding in spoken language models. Vox-Infinity systematically extends audio history along two dimensions: turn count and turn duration. It covers a diverse range of representative scenarios with varying interaction structures and semantic complexity. Crucially, Vox-Infinity provides explicit answer-provenance annotations and organizes samples according to the amount of historical context required to resolve each query, enabling precise and length-aware evaluation. Extensive evaluations of seven representative spoken language models reveal a clear overall recency effect: models generally achieve higher accuracy when answer-supporting evidence is closer to the query, but struggle to retrieve and use evidence located farther back in the dialogue history. Cases and datasets are available at https://vox-infinity.github.io.

↑