arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过高效的段级到视频级监督增强长视频理解的局部推理能力

Enhancing Localized Reasoning for Long Video Understanding via Efficient Segment-to-Video Supervision

Beibei Zhang, Chao Xu, Jun Lan, Zongyi Li, Lai Wei, Huijia Zhu, Tongwei Ren

arXiv 2608.20814首次发表:更新:

发表机构

Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长视频理解中MLLMs易受干扰噪声影响的问题,提出S2V方法,通过段级VQA训练模型,提升了LVU性能与训练推理效率。

AI 中文摘要

尽管多模态大语言模型(MLLMs)在视频理解方面展现出巨大潜力,但长视频理解(LVU)仍面临挑战,因为复杂冗长上下文里的干扰噪声会掩盖局部细节,误导MLLMs生成错误答案。近期研究通过激励深度推理纳入相关证据来缓解这些问题,但这些方法存在两个主要缺陷:其一,它们采用的强化微调框架(RFT)会产生大量训练开销,包括高昂的标注成本和复杂的奖励设计;其二,部分方法中的自反思与迭代感知机制会导致输出冗长且推理延迟高。为缓解这些问题,我们提出一种新颖的段级到视频级监督方法(S2V),以高效增强LVU中的细粒度推理能力。具体而言,我们基于局部段生成问答对(VQA),随后将这些基于段的VQA回传到整个视频用于训练。由于聚焦于短段,基于段的VQA自然能注意到从全视频视角易被忽略的细节,在这类数据上训练可强制MLLMs将细粒度细节与问答正确关联,同时规避全视频中的干扰噪声。S2V训练仅涉及基于简单准确率奖励的强化学习(RL),且仅使用10K个VQA样本;最终的S2V模型通过单次前向传播并使用有限输出令牌预测答案。实验结果表明,S2V在多个LVU基准上均能持续提升LVU性能,不仅在LVU准确率上优于通用MLLMs和基于推理的方法,在训练与推理效率上也表现更优。

英文摘要

Though Multimodal Large Language Models (MLLMs) have shown impressive potential in video understanding, long video understanding (LVU) remains challenging since distracting noise in complex and lengthy contexts can obscure localized details, misleading MLLMs to produce incorrect answers. Recent works mitigate these issues by incentivizing deep reasoning to include relevant evidence. However, these methods have two main problems: First, the reinforcement fine-tuning framework (RFT) they leveraged incurs substantial training overheads, including high annotation costs and complicated reward designs. Second, the self-reflective and iterative-perception mechanism in some methods causes lengthy outputs and high inference latency. To alleviate these problems, we propose a novel Segment-to-Video Supervision} method (S2V) to efficiently enhance fine-grained reasoning in LVU. Specifically, we generate question answer pairs (VQA) based on localized segments, and then transfer these segment-based VQA back to the whole video for training. Due to focusing on short segments, segment-based VQA can naturally notice details which tend to be overlooked from a whole-video perspective. Training on such data can enforce MLLMs to correctly associate fine-grained details with QA while avoiding distracting noise in the whole video. The S2V training involves just reinforcement learning (RL) with a simple accuracy reward based on only 10K VQA samples and the resulting S2V model predicts answer using a single forward pass with limited output tokens. Experimental results demonstrate that S2V can consistently improve LVU performance across multiple LVU benchmarks, outperforming both general MLLMs and reasoning-based methods not only in LVU accuracy but also in training and inference efficiency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑