Video-R4:通过视觉沉思强化文本丰富的视频推理
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
- University of Rochester(罗切斯特大学)
- Sony Group Corporation(索尼集团)
- MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Video-R4通过视觉沉思机制提升文本丰富视频的推理能力,采用多阶段学习框架实现像素基础的多模态推理。
AI中文摘要:
理解包含大量文本的视频需要阅读短暂的文本提示,这些提示往往需要反复检查。然而,大多数视频问答模型依赖于对固定帧的单次感知,导致幻觉和在细粒度证据上的失败。受人类暂停、放大和重新阅读关键区域的启发,我们引入了Video-R4(通过视觉沉思强化文本丰富的视频推理),一种视频推理LMM,执行视觉沉思:迭代选择帧、放大信息区域、重新编码检索的像素并更新其推理状态。我们构建了两个具有可执行沉思轨迹的数据集:Video-R4-CoT-17k用于监督练习,Video-R4-RL-30k用于强化学习。我们提出了一种多阶段沉思学习框架,逐步微调7B LMM,通过SFT和基于GRPO的RL学习原子和混合的视觉操作。Video-R4-7B在M4-ViteVQA上取得了最先进的结果,并进一步扩展到多页文档问答、幻灯片问答和通用视频问答,证明了迭代沉思是像素基础多模态推理的有效范式。项目页面:https://yunlong10.github.io/Video-R4/
英文摘要:
Understanding text-rich videos requires reading small, transient textual cues that often demand repeated inspection. Yet most video QA models rely on single-pass perception over fixed frames, leading to hallucinations and failures on fine-grained evidence. Inspired by how humans pause, zoom, and re-read critical regions, we introduce Video-R4 (Reinforcing Text-Rich Video Reasoning with Visual Rumination), a video reasoning LMM that performs visual rumination: iteratively selecting frames, zooming into informative regions, re-encoding retrieved pixels, and updating its reasoning state. We construct two datasets with executable rumination trajectories: Video-R4-CoT-17k for supervised practice and Video-R4-RL-30k for reinforcement learning. We propose a multi-stage rumination learning framework that progressively finetunes a 7B LMM to learn atomic and mixing visual operations via SFT and GRPO-based RL. Video-R4-7B achieves state-of-the-art results on M4-ViteVQA and further generalizes to multi-page document QA, slides QA, and generic video QA, demonstrating that iterative rumination is an effective paradigm for pixel-grounded multimodal reasoning. Project Page: https://yunlong10.github.io/Video-R4/