通过强大的音频感知增强长格式全模态理解
Empowering Long-form Omni-modal Understanding with Robust Audio Perception
浏览论文内容
中文总结 AI 辅助
为解决全模态理解不足问题,提出AVDC数据集及AVDC-QA-CoT数据集,利用现成模型标注视频,采用两阶段训练范式,在多下游任务实验中取得显著性能提升,推动全模态感知发展。
中文摘要 AI 辅助
大规模多模态模型在视觉语言任务上取得显著进展,但全模态理解仍未充分探索,主要因缺乏含丰富对齐音频线索的数据集。为此提出AVDC数据集,利用现成模型用三方字幕标注视频,明确捕捉模态细微差别和跨模态交互。在此基础上引入AVDC-QA-CoT数据集促进视听推理。采用两阶段训练范式,在不同下游任务实验中均取得显著性能提升。
英文摘要
Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich, explicitly aligned auditory cues. To bridgethis gap, we present AVDC (Audio-Visual Decoupled Captions), a large-scaledataset designed to disentangle visual and auditory semantics. Specifi-cally, we propose an automated pipeline that leverages off-the-shelf mod-els to annotate videos with tripartite captions: visual-only (V), audio-only (A), and joint audio-visual (AV). This decoupled structure explic-itly captures both modality-specific nuances and complex cross-modalinteractions. Building upon this, we introduce AVDC-QA-CoT, a Chain-of-Thought augmented question-answering dataset to foster audio-visualreasoning. To fully exploit these resources, we employ a two-stage train-ing paradigm: omni-modal caption generation pre-training on AVDC, fol-lowed by instruction tuning on AVDC-QA-CoT. Extensive experiments acrossdiverse downstream tasks, spanning video captioning, audio-centric anal-ysis, and omni-modal benchmarks, demonstrate consistent and signifi-cant performance gains, showing the efficacy of our proposed datasetsand training strategy in advancing omni-modal perception. Code anddataset are related on https://radiant0726.github.io/AVDC-web/.
发表机构
- SAI, Shanghai Jiao Tong University(上海交通大学 上海人工智能研究院)
- Zhejiang University(浙江大学)
- Shanghai AI Lab(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。