arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10299cs.LG

通过强大的音频感知增强长格式全模态理解

Empowering Long-form Omni-modal Understanding with Robust Audio Perception

Kaiying Yan, Luoyi Sun, Xiao Zhou, Weidi Xie

首次发表
浏览论文内容

中文总结 AI 辅助

为解决全模态理解不足问题,提出AVDC数据集及AVDC-QA-CoT数据集,利用现成模型标注视频,采用两阶段训练范式,在多下游任务实验中取得显著性能提升,推动全模态感知发展。

中文摘要 AI 辅助

大规模多模态模型在视觉语言任务上取得显著进展,但全模态理解仍未充分探索,主要因缺乏含丰富对齐音频线索的数据集。为此提出AVDC数据集,利用现成模型用三方字幕标注视频,明确捕捉模态细微差别和跨模态交互。在此基础上引入AVDC-QA-CoT数据集促进视听推理。采用两阶段训练范式,在不同下游任务实验中均取得显著性能提升。

英文摘要

Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich, explicitly aligned auditory cues. To bridgethis gap, we present AVDC (Audio-Visual Decoupled Captions), a large-scaledataset designed to disentangle visual and auditory semantics. Specifi-cally, we propose an automated pipeline that leverages off-the-shelf mod-els to annotate videos with tripartite captions: visual-only (V), audio-only (A), and joint audio-visual (AV). This decoupled structure explic-itly captures both modality-specific nuances and complex cross-modalinteractions. Building upon this, we introduce AVDC-QA-CoT, a Chain-of-Thought augmented question-answering dataset to foster audio-visualreasoning. To fully exploit these resources, we employ a two-stage train-ing paradigm: omni-modal caption generation pre-training on AVDC, fol-lowed by instruction tuning on AVDC-QA-CoT. Extensive experiments acrossdiverse downstream tasks, spanning video captioning, audio-centric anal-ysis, and omni-modal benchmarks, demonstrate consistent and signifi-cant performance gains, showing the efficacy of our proposed datasetsand training strategy in advancing omni-modal perception. Code anddataset are related on https://radiant0726.github.io/AVDC-web/.

发表机构

  • SAI, Shanghai Jiao Tong University(上海交通大学 上海人工智能研究院)
  • Zhejiang University(浙江大学)
  • Shanghai AI Lab(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑