arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02940cs.CLcs.AI

聆听隐态:基于隐态交互的大型音频语言模型自校正语音识别

Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions

发表机构英伟达 · 卡内基梅隆大学 · 台湾大学
查看机构详情
  • NVIDIA(英伟达)
  • Carnegie Mellon University(卡内基梅隆大学)
  • National Taiwan University(台湾大学)

机构由 AI 辅助整理,请以论文原文为准。

Chan-Jan Hsu, Jaeyeon Kim, Chao-Han Huck Yang, Shinji Watanabe, Hung-yi Lee, Carlos Busso

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对现有 ASR 系统整合 LLM 的两种策略结合不足的问题,提出 Hybrid Search 针对性校正策略,利用 LoRA 适配场景下的基础 LLM 隐态交互特征优化目标 token,显著提升了 ASR 推理性能。

中文摘要 AI 辅助

近年来,自动语音识别(ASR)系统越来越多地整合大型语言模型(LLM)以利用其语义知识,方式要么是通过 logit 融合的外部整合,要么是通过 warm 初始化的内部整合。然而,如何有效结合这两种策略仍未得到充分探索。在本研究中,我们通过利用 warm 初始化 LLM 型 ASR 模型自身的预适配基础 LLM 来优化该模型,重点关注基础 LLM 被保留的 LoRA 适配场景。为实现这一目标,我们提出了 Hybrid Search,这是一种受两个观察结果启发的针对性校正策略。第一,表征 LLM 型 ASR 隐态与基础 LLM 隐态之间关系的交互特征,提供了关于 token 语义依赖程度的有用信号。第二,选择性优化具有高语义依赖度的目标 token,能大幅超越包括重评分和后期融合在内的朴素全局 LLM 校正方法,显著提升 ASR 性能。我们的分析表明,即使在通过 warm 初始化完成语义知识迁移后,LLM 型 ASR 模型仍可利用其基础 LLM 进一步提升推理时的性能。

英文摘要

Recent automatic speech recognition (ASR) systems increasingly integrate large language models (LLMs) to leverage their semantic knowledge, either externally through logit fusion or internally through warm initialization. However, how to effectively combine these two strategies remains underexplored. In this work, we refine warm-initialized LLM-based ASR models by leveraging their own pre-adaptation base LLMs, focusing on LoRA-adapted settings where the base LLM is preserved. To achieve this, we propose Hybrid Search, a targeted correction strategy motivated by two observations. First, interaction features that characterize the relationship between LLM-based ASR hidden states and base-LLM hidden states provide informative signals about a token's degree of semantic dependence. Second, selectively refining targeted tokens with high semantic dependence improves ASR performance far beyond naive global LLM-correction methods including rescoring and late fusion. Our analysis suggests that, even after semantic knowledge transfer through warm initialization, LLM-based ASR models can still leverage their base LLM to further improve inference-time performance.

补充信息

↑