arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32536cs.SDcs.AIcs.CLcs.MMeess.AS

音频大语言模型在行动前会“听”吗?诊断语音代理中的声学上下文门控

Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents

Yanjie Zhang, Nanchen Hu, Yushi Sun

中文总结 AI 辅助

本文提出VGBench基准,诊断语音代理在声学上下文中的门控能力,发现原始模型常不弃权,而VoxGate监督训练显著提升切换静默率至91.3%。

中文摘要 AI 辅助

音频语言模型能够识别语音指令并调用工具,但代理必须首先决定声学和对话上下文是否值得采取行动。我们引入了VGBench,一个包含1,018个条目的诊断基准,用于评估侧向交谈、自言自语和说话者切换场景下的动作级寻址性。每个条目使用一个共享的动作空间,包括静默、工具调用和自然语言回答。说话者切换对在保持指定词语不变的同时,通过声源、距离渲染和时间边界来定义受控的佩戴者到旁观者的转换。六个原始音频大语言模型和三种免训练适配方法通常能识别目标工具,但在这种转换下很少抑制行动;最高的原始切换静默率为14%。然后,我们使用VoxGate作为训练后的案例研究。监督训练使91.3%的切换指令静默,同时为所有附近的佩戴者指令和纯文本控制选择正确的工具。一个探索性的GRPO阶段具有类似的切换性能;侧向交谈准确率从68.4%提升到70.9%,自言自语静默率从52.0%提升到60.0%。因子化控制识别出独立的声源变化效应,而对远场操控的敏感性因声学渲染而异。因此,该基准衡量的是多线索声学上下文门控,而非孤立的说话者身份。

英文摘要

Audio language models can recognize spoken commands and invoke tools, but an agent must first decide whether the acoustic and conversational context warrants action. We introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across side-talk, self-talk, and speaker-switch scenarios. Each item uses a shared action space comprising silence, a tool call, and a natural-language answer. Speaker-switch pairs hold the specified words fixed while source, distance rendering, and a temporal boundary define a controlled wearer-to-bystander shift. Six raw Audio LLMs and three training-free adaptations often identify the target tool yet rarely withhold action under this shift; the highest raw switch mute rate is 14%. We then use VoxGate as a post-training case study. Supervised training mutes 91.3% of switched commands while choosing the correct tool for all nearby wearer commands and text-only controls. An exploratory GRPO stage has similar switch performance; side-talk accuracy rises from 68.4% to 70.9%, and self-talk muting from 52.0% to 60.0%. Factorized controls identify an independent source-change effect, while sensitivity to the far-field manipulation varies across acoustic renderings. The benchmark therefore measures multi-cue acoustic-context gating rather than isolated speaker identity.

↑