超越熵:面向视频推理的自诊断多角色令牌优化
Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning
- City University of Hong Kong(香港城市大学)
- Institute of Information Engineering, Chinese Academy of Sciences (CAS)(中国科学院信息工程研究所)
- Beihang University(北京航空航天大学)
- AMAP, Alibaba Group(阿里巴巴集团高德地图)
- Computer Science, Beijing Institute of Technology(北京理工大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出DyCPO协同演化框架,通过多角色依赖度量与动态反事实信号,解决视频推理中令牌级信用分配模糊问题,在基准上实现持续性能提升。
AI中文摘要:
基于可验证奖励的强化学习已显著推进了多模态推理,但其仍从根本上受限于模糊的令牌级信用分配。虽然高熵令牌启发式方法鼓励可能性探索,但将其直接扩展至视频推理往往会导致冗长的推理过程,因为模型过度依赖高熵视觉激活。依赖基于反事实的视觉令牌定位进行信用分配的替代方法,也倾向于过度优先进行视觉探索,而牺牲了用于答案推导的关键推理线索,从而加剧了无关视觉细节的干扰。此外,这些方法采用静态反事实策略,无法在训练过程中与策略共同演化。本文提出DyCPO,一种协同演化框架,联合优化可靠的令牌选择与自适应反事实干预。该框架构建多角色依赖度量,在令牌级对比学习中平衡视觉探索与答案相关性挖掘,同时抑制仅用于探索的填充令牌和虚假视觉噪声。DyCPO不依赖静态反事实先验,而是从模型自身的成功与失败轨迹中动态推导反事实信号,实现自诊断分析以及优化目标与策略的协同演化。在复杂视频推理和通用视频理解基准上的大量实验表明,性能持续提升,确立了DyCPO作为多模态强化学习中稳健的令牌级信用分配范式。
英文摘要:
Reinforcement learning with verifiable rewards has substantially advanced multimodal reasoning, yet it remains fundamentally limited by ambiguous token-level credit assignment. While high-entropy token heuristics encourage possibility exploration, naively extending them to video reasoning tends to induce lengthy reasoning, as the model becomes overly reliant on high-entropy visual activations. Alternative approaches that rely on counterfactual-based visual token localization for credit assignment also tend to over-prioritize visual exploration at the expense of decisive reasoning cues for answer derivation, thereby exacerbating the interference from spurious visual nuances. Moreover, these methods employ static counterfactual strategies that fail to co-evolve with the policy during training. In this paper, we introduce DyCPO, a co-evolutionary framework that jointly optimizes reliable token selection and adaptive counterfactual intervention. It constructs a multi-role dependence metric to balance visual exploration and answer-relevance mining in token-wise contrastive learning, while suppressing exploration-only filler tokens and spurious visual noise. Rather than relying on static counterfactual priors, DyCPO dynamically derives counterfactual signals from the model's own successful and failed rollouts, enabling self-diagnostic analysis and co-evolution of the optimization objective with the policy. Extensive experiments on complex video reasoning and general video understanding benchmarks demonstrate consistent performance improvements, establishing DyCPO as a robust token-level credit assignment paradigm for multimodal reinforcement learning.