arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26099cs.CV

异常视频理解的测试时强化学习

Test-time Reinforcement Learning for Anomalous Video Understanding

Huining Li, Yuxiang Duan, Jiyang Tan, Qian Li, MingCai Chen, Jian Zhang, Xingdong Sheng, Yuntao Du

首次发表
浏览论文内容

中文总结 AI 辅助

针对异常视频理解中视频大模型适应不足的问题,提出测试时强化学习框架,通过双查询过滤、熵感知奖励和虚拟负锚点机制,在VAU-Bench上显著提升准确率至90%。

中文摘要 AI 辅助

异常视频理解旨在识别视频中的异常事件,并解释其语义含义,而不仅仅是简单的异常检测。近期,视频大语言模型(Video-LLMs)已展现出在该任务上具有前景的零样本能力,然而,由于对多样化的异常模式和不断变化的环境适应不足,其性能仍然受限。测试时强化学习提供了一种有前景的解决方案,使模型能够通过自生成的反馈信号进行改进,而无需额外的人工标注。然而,将其应用于异常视频理解仍面临三个挑战:(1)当共识较弱时,生成的伪标签可能不可靠;(2)二元奖励设计无法捕捉模型生成中的不确定性,导致优化信号无效;(3)一致的行动组获得相同的奖励,导致组相对优势崩溃,消除了有效的策略梯度信号。为解决这些挑战,我们提出了一种新颖的用于异常视频理解的测试时强化学习框架,引入了双查询一致性过滤、熵感知共识奖励和虚拟负锚点机制。该框架通过语义等价查询之间的一致性保留可靠样本,将答案一致性与生成不确定性相结合进行奖励估计,并引入虚拟负锚点在一致行动组中创造奖励变化,从而保留有效的组相对优化信号。在VAU-Bench上的实验表明,我们的方法优于所比较的冻结基线和监督基线。在VAU-Bench的ECVA子集上,增益最为显著,在思考模式下,准确率从75.81%提升至90.00%,相对于冻结骨干网络。

英文摘要

Anomalous video understanding aims to identify abnormal events in videos and interpret their semantic meanings beyond simple anomaly detection. Recent video large language models (Video-LLMs) have demonstrated promising zero-shot capabilities for this task, yet their performance remains limited due to insufficient adaptation to diverse anomaly patterns and evolving environments. Test-time reinforcement learning offers a promising solution by enabling models to improve through self-generated feedback signals without requiring additional human annotations. However, applying it to anomalous video understanding remains challenging due to three issues: (1) generated pseudo-labels can be unreliable when consensus is weak; (2) binary reward designs fail to capture uncertainty in model generations, resulting in ineffective optimization signals; and (3) unanimous rollout groups receive identical rewards, causing group-relative advantages to collapse and eliminating effective policy-gradient signals. To address these challenges, we present a novel test-time reinforcement learning framework for anomalous video understanding by introducing dual-query consistency filtering, an entropy-aware consensus reward, and a virtual negative anchor mechanism. The framework retains reliable samples through consistency across semantically equivalent queries, combines answer agreement with generation uncertainty for reward estimation, and introduces a virtual negative anchor to create reward variation in unanimous rollout groups, thereby preserving effective group-relative optimization signals. Experiments on VAU-Bench show that our method outperforms the compared frozen and supervised baselines. The gains are most pronounced on the ECVA subset of VAU-Bench with thinking, where accuracy improves from 75.81% to 90.00% relative to the frozen backbone.

↑