arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30625cs.AI

音频大语言模型知道自己何时听不到你

Audio LLMs Know When They Can't Hear You

发表机构苹果公司 · 加州大学圣地亚哥分校
查看机构详情
  • Apple(苹果公司)
  • University of California San Diego(加州大学圣地亚哥分校)

机构由 AI 辅助整理,请以论文原文为准。

Amirhosein Javadi, Richa Dixit, Mehrdad Farajtabar, Minsik Cho, Devang Naik, Mohammad Samragh

首次发表
浏览论文内容

中文总结 AI 辅助

本文研究音频大语言模型能否识别自身转录的不可靠性,发现其自评能力差,但音频编码器表示可强预测可靠性,据此提出轻量级预测器,域内和跨域宏F1分别达81.10%和78.09%,优于基线。

中文摘要 AI 辅助

音频大语言模型允许用户通过语音与模型交互。当输入录音质量过于退化时,模型可能会误解用户的查询,并基于错误的转录做出响应。在本文中,我们研究模型条件转录可靠性:即音频大语言模型能否识别其自身转录的不可靠性。我们首先提示音频大语言模型评估其自身转录是否可靠,并发现该模型对自身转录可靠性的判断能力较差:在大多数情况下,它预测其转录将是可靠的。我们发现现有方法,包括语音质量预测器、音频大语言模型生成不确定性和转录条件下的词错误率估计,为检测转录失败提供的信号有限。相比之下,我们发现转录可靠性在模型的音频编码器表示中得到了强烈体现。基于这一观察,我们设计了一个轻量级可靠性预测器,该预测器在冻结音频编码器提取的表示上运行,并在生成之前预测可靠性类别。当预测用户的语音查询不可靠时,该可靠性预测器可以触发用户的澄清请求,同时允许可靠的查询继续进行,而无需修改底层音频大语言模型。我们的预测器实现了81.10%的域内和78.09%的跨域宏F1分数,分别比最强基线高出10.33和11.93个百分点。最后,我们表明可靠性标签可以在音频大语言模型家族之间转移,并且转移性能与其模型特定可靠性边界的对齐密切相关。

英文摘要

Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable. We first prompt the Audio LLM to assess whether its own transcription would be reliable, and find that the model is a poor judge of its own transcription reliability: in most cases, it predicts that its transcription will be reliable. We find that existing approaches, including speech quality predictors, audio LLM generation uncertainty, and transcript-conditioned WER estimation, provide limited signals for detecting transcription failures. In contrast, we discover that transcription reliability is strongly represented in the model's audio-encoder representations. Based on this observation, we devise a lightweight reliability predictor that operates on representations extracted by the frozen audio encoder and predicts the reliability class before generation. The reliability predictor can trigger a clarification request from the user when their voice query is predicted to be unreliable, while allowing reliable queries to proceed without modifying the underlying Audio LLM. Our predictor achieves 81.10% in-domain and 78.09% cross-domain macro-F1 scores, outperforming the strongest baselines by 10.33 and 11.93 points, respectively. Finally, we show that reliability labels can transfer across Audio LLM families, and that transfer performance is closely related to the alignment of their model-specific reliability boundaries.

补充信息

↑