arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁在说什么:音视频大语言模型中的符号化三模态绑定机制

Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs

Jihoo Jung, Youngjoon Jang, Joon Son Chung

arXiv 2609.31193首次发表:更新:

发表机构

KAIST; VGG, University of Oxford(韩国科学技术院; 牛津大学视觉几何组)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多说话人视频中“谁在说什么”的推理难题,本文揭示AVLLMs通过符号变量实现三模态绑定,并发现失败源于音视频错位,提出基于主动说话人检测的免训练提示方法,在多个基准上取得即时效能提升。

AI 中文摘要

当前音视频大语言模型(AVLLMs)在处理包含多说话人对话的视频推理任务时面临困难。在这类视频中,解析“谁在说什么”至关重要,这需要三模态(文本-音频-视觉)绑定。受这些挑战的驱动,我们系统地研究了AVLLMs中如何实现这种三模态绑定。具体而言,我们识别出AVLLMs中利用模态特定符号变量的涌现式符号化三模态绑定机制。通过将音频和视觉成分编码为符号变量——分别捕获时间上的话语序列和空间上的实体坐标——模型在此抽象空间内建立跨模态链接。关键的是,我们揭示当三模态绑定失败时,其崩溃主要源于音频-视觉连接的对齐错误。为克服这一瓶颈,我们引入一种利用现成主动说话人检测(ASD)模型的音视频提示方法。通过简单地在活跃说话人上叠加视觉边界框,这种免训练方法在四个以对话为中心的基准上立即带来性能提升。此外,在这些ASD提示视频上进行少于300步的轻量级微调,可将这些增益扩展到三个通用音视频基准,表明我们方法的泛化性。

英文摘要

Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal binding is achieved in AVLLMs. Specifically, we identify emergent symbolic trimodal binding mechanisms in AVLLMs that utilize modality-specific symbolic variables. By encoding auditory and visual components into symbolic variables-capturing temporal utterance sequences and spatial entity coordinates, respectively-the model establishes cross-modal linking within this abstract space. Crucially, we reveal that when trimodal binding fails, the breakdown predominantly stems from misaligned audio-visual connections. To overcome this bottleneck, we introduce an audio-visual prompting method utilizing an off-the-shelf Active Speaker Detection (ASD) model. By simply overlaying visual bounding boxes on active speakers, this training-free approach yields immediate performance gains across four conversation-centric benchmarks. Moreover, lightweight fine-tuning of fewer than 300 steps on these ASD-prompted-videos extends these gains to three general AV benchmarks, suggesting the generalizability of our method.

CommentsAccepted by NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑