无训练的以语音为中心的全模态理解:基于冻结的视觉语言模型
Training-Free Speech-Centric Omni Understanding with Frozen VLMs
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出即插即用框架TFO,通过Whisper和VLM现有语言接口将冻结VLM转为以语音为中心的全模态模型,在56个基准、21种语言上的对比中表现具竞争力,保留了VLM原有能力且成本更低。
AI中文摘要:
视听理解仍然具有挑战性,因为模型必须联合解释语音内容、视觉事件及其时间关系。现有的全模态模型通常引入专用音频编码器,并依赖昂贵的音-视-文本训练,这将全模态能力与特定的视觉语言模型(VLM)主干紧密耦合,可能会削弱其现有的视觉和推理能力。这引发了三个问题:是否每个新的VLM都需要原生全模态训练,是否可以在保留原始主干的同时添加以语音为中心的全模态能力,以及更丰富的声学表征在哪些场景下仍然必不可少。我们提出了无训练全模态(TFO),这是一种即插即用框架,无需架构修改或多模态重新对齐,即可将冻结的VLM转换为以语音为中心的全模态模型。TFO使用Whisper提取经过置信度过滤的带时间戳的文本转录,并通过VLM现有的语言接口进行路由,同时保持其视觉通路不变。在与原生全模态模型在56个基准和21种语言上的匹配对比中,TFO在视听理解方面具有竞争力,在所有5种模型设置下平均仅音频性能均有所提升,并在多语言语音任务中实现了显著增益。与相应的原生全模态检查点相比,冻结VLM还通常能保留更强的图像/视频理解、视觉 grounding、编码、数学推理和医学问答能力。这些结果表明,强大的以语音为中心的全模态理解通常可以通过模块化的音频到语言路由而非昂贵的主干特定训练来实现。
英文摘要:
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual events, and their temporal relationships. Existing omni models typically introduce dedicated audio encoders and rely on expensive audio-video-text training, tightly coupling omni capability to specific VLM backbones and potentially weakening their existing visual and reasoning abilities. This raises three questions: whether native omni training is necessary for every new VLM, whether speech-centric omni capability can be added while preserving the original backbone, and where richer acoustic representations remain essential. We introduce Training-Free Omni (TFO), a plug-and-play framework that converts a frozen VLM into a speech-centric omni model without architectural modification, or multimodal re-alignment. TFO uses Whisper to extract confidence-filtered, timestamped transcripts and routes them through the VLM's existing language interface, while leaving its visual pathway unchanged. Across matched comparisons with native omni models on 56 benchmarks and 21 languages, TFO is competitive on audio-visual understanding, improves average audio-only performance across all five model settings, and achieves substantial multilingual speech gains. Freezing the VLM also generally preserves stronger image/video understanding, visual grounding, coding, mathematical reasoning, and medical question answering than the corresponding native omni checkpoints. These results show that strong speech-centric omni understanding can often be obtained through modular audio-to-language routing rather than costly backbone-specific training.