Video-KTR: 通过关键令牌 attribution 增强视频推理
Video-KTR: Reinforcing Video Reasoning via Key Token Attribution
- ByteDance(字节跳动)
- School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电子与电气工程学院)
- National University of Singapore(国立新加坡大学)
- College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Video-KTR通过结合三种attribution信号,提升视频推理的准确性和可解释性,实现状态-of-the-art性能。
AI中文摘要:
强化学习(RL)在增强多模态大语言模型的推理能力方面显示出强大潜力,但现有视频推理方法往往依赖于粗粒度的序列级奖励或单一因素的令牌选择,忽略了视觉输入、时间动态和语言输出之间精细的联系,限制了准确性和可解释性。我们提出了Video-KTR,一种模态感知的策略塑造框架,通过结合三种 attribution 信号进行选择性、令牌级的RL:(1)通过反事实遮蔽识别出的视觉感知令牌,以揭示感知依赖性;(2)通过帧洗牌检测出的时间感知令牌,以暴露时间敏感性;(3)高熵令牌信号预测不确定性。通过仅强化这些关键令牌,Video-KTR 将学习集中在语义信息丰富、模态敏感的内容上,同时过滤掉低价值的令牌。在五个具有挑战性的基准测试中,Video-KTR 获得了最先进的或高度竞争性的结果,其在Video-Holmes上达到了42.7%(超越GPT-4o),并在推理和通用视频理解任务上实现了持续的提升。消融研究验证了 attribution 信号的互补作用以及目标令牌级更新的鲁棒性。总体而言,Video-KTR 提高了准确性和可解释性,提供了一种简单的、可直接使用的RL扩展方法,用于复杂的视频推理。我们的代码和模型可在https://github.com/zywang0104/Video-KTR上获得。
英文摘要:
Reinforcement learning (RL) has shown strong potential for enhancing reasoning in multimodal large language models, yet existing video reasoning methods often rely on coarse sequence-level rewards or single-factor token selection, neglecting fine-grained links among visual inputs, temporal dynamics, and linguistic outputs, limiting both accuracy and interpretability. We propose Video-KTR, a modality-aware policy shaping framework that performs selective, token-level RL by combining three attribution signals: (1) visual-aware tokens identified via counterfactual masking to reveal perceptual dependence; (2) temporal-aware tokens detected through frame shuffling to expose temporal sensitivity; and (3) high-entropy tokens signaling predictive uncertainty. By reinforcing only these key tokens, Video-KTR focuses learning on semantically informative, modality-sensitive content while filtering out low-value tokens. Across five challenging benchmarks, Video-KTR achieves state-of-the-art or highly competitive results, achieving 42.7\% on Video-Holmes (surpassing GPT-4o) with consistent gains on both reasoning and general video understanding tasks. Ablation studies verify the complementary roles of the attribution signals and the robustness of targeted token-level updates. Overall, Video-KTR improves accuracy and interpretability, offering a simple, drop-in extension to RL for complex video reasoning. Our code and models are available at https://github.com/zywang0104/Video-KTR.