arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34223cs.CVcs.SD

揭示音视频大语言模型中的序数匹配偏差

Uncovering Ordinal-Matching Bias in Audio-Visual LLMs

发表机构韩国科学技术院 · 牛津大学
查看机构详情
  • Korea Advanced Institute of Science and Technology(韩国科学技术院)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Jihoo Jung, Youngjoon Jang, Hyebin Cho, Suho Yoo, Joon Son Chung

首次发表
浏览论文内容

中文总结 AI 辅助

针对多说话人场景中音视频大模型将语音与说话人错误关联的问题,本文揭示了序数匹配偏差,并提出序数解耦微调(OD-FT)方法,通过随机化合成视频微调,有效抑制偏差并提升真实视频理解性能。

中文摘要 AI 辅助

这项工作旨在改进音视频大语言模型(AVLLMs)在多说话人场景中将语音与正确的可见说话人关联的能力。我们发现当前的AVLLMs在此任务上频繁失败,并分析了这些失败的本质。为此,我们构建了一个合成诊断数据集,其中多个可见说话人各自说出一个单词。对该语料库的分析揭示了三个近期开源AVLLMs中存在一致的错误模式:模型通过简单地将口语句子的顺序与可见面孔从左到右、从上到下的排列相匹配来归因话语,而不是依赖音频-视觉线索(如唇部同步)。我们将这种行为称为“序数匹配偏差”。我们进一步表明,这种偏差可以通过一个简单的补救措施——序数解耦微调(OD-FT)——得到显著缓解,该措施在空间位置和说话顺序独立随机化的合成视频上对模型进行微调。尽管仅使用400个合成训练视频,OD-FT不仅抑制了序数匹配偏差,还提高了模型在真实世界视频上的音视频理解能力,在三个音视频基准上,Qwen2.5-Omni平均提升8.27%,video-SALMONN2+平均提升2.57%。

英文摘要

This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic dataset in which multiple visible speakers each utter a single word. Analysis on this corpus reveals a consistent error pattern across three recent open-source AVLLMs: models attribute utterances by simply matching the order of spoken sentences with the left-to-right, top-to-bottom arrangement of visible faces, rather than relying on audio-visual cues such as lip synchronization. We term this behavior \emph{ordinal-matching bias}. We further show that this bias can be substantially mitigated through a simple remedy, Ordinal-Decoupled Fine-Tuning (OD-FT), in which models are fine-tuned on synthetic videos where spatial positions of speakers and speaking order are independently randomized. Despite using only 400 synthetic training videos, OD-FT not only suppresses ordinal-matching bias but also improves audio-visual understanding on real-world videos, yielding average gains of 8.27\% for Qwen2.5-Omni and 2.57\% for video-SALMONN2+ across three audio-visual benchmarks.

补充信息

↑