发表机构
Meta Reality Labs(Meta现实实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对音频大模型被动响应的问题,提出打断与静默建模(ISM)范式,通过特殊标记实现主动音频辅助,在ESC-50和Epic-Sounds上取得最优性能,平均延迟3.5秒。
AI 中文摘要
音频大语言模型(AudioLLMs)以被动方式运行,仅在收到查询时做出响应。我们引入了主动音频辅助的概念,即 AudioLLM 监控音频流,并根据单一自然语言意图自主决定何时提醒用户,这一概念受到面向聋人和听障人士的可穿戴应用的启发。我们提出了打断与静默建模(ISM),这是一种与模型无关的范式,通过两个特殊标记 \ exttt{<interrupt>} 和 \ exttt{<silent>} 将主动决策嵌入到 LLM 解码中,捕捉四种状态:起始检测、持续相关触发、无关抑制和去重。将 ISM 应用于 Qwen2-Audio-7B,在 ESC-50 上实现了 99.6% 的打断 F1 分数和完美的去重召回率。在嘈杂的 Epic-Sounds 厨房音频上,ISM 在没有领域特定训练的情况下实现了最高的打断 F1 分数,是唯一在不过度触发或过度抑制的情况下保持强起始检测的方法。流式评估证实了其实时可行性,平均延迟为 3.5 秒。
英文摘要
Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM monitors an audio stream and autonomously decides when to alert the user from a single natural-language intent, motivated by wearable applications for Deaf and Hard of Hearing users. We propose Interrupt and Silent Modeling (ISM), a model-agnostic paradigm that embeds proactive decisions into LLM decoding via two special tokens: \texttt{<interrupt>} and \texttt{<silent>}, capturing four states: onset detection, sustained-relevance triggering, irrelevance suppression, and de-duplication. Applied to Qwen2-Audio-7B, ISM achieves 99.6\% interrupt F1 and perfect de-duplication recall on ESC-50. On noisy Epic-Sounds kitchen audio, ISM achieves the highest interrupt F1 without domain-specific training, the only method maintaining strong onset detection without over-triggering or over-suppression. Streaming evaluation confirms real-time viability with 3.5-second average latency.
CommentsAccepted at Interspeech 2026