发表机构
AMAP, Alibaba Group(阿里巴巴集团AMAP)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频时序定位中视觉证据与时间戳错位的问题,提出CAVE方法,通过边界证据奖励、轻量级热身及性能感知门控,在公开基准上验证了有效性。
AI 中文摘要
大型视觉语言模型(LVLMs)通过强化学习(RL)在视频时序定位(VTG)任务中取得了显著的性能提升。然而,现有方法主要依赖仅评估最终预测区间的结果正确性奖励,对与边界相关的视觉证据及其与时间戳预测的对应关系约束不足。在本文中,我们深入研究时间戳预测及其底层的边界级视觉证据,发现广泛使用的基准中普遍存在视觉证据与预测时间戳之间的错位问题。为解决该问题,我们提出了能力感知视觉边界证据对齐(CAVE)方法,该方法通过边界特定的视觉证据奖励增强定位优化,以缓解证据与时间戳的错位。具体而言,为明确表示边界特定的视觉证据,CAVE引入边界特定的证据标记,并通过轻量级监督热身初始化其结构化生成和不同的边界语义。在RL过程中,视觉边界证据对齐奖励强化了真实边界内特殊证据标记的视觉注意力,从而促进视觉证据与时间边界的对齐。此外,我们设计了面向证据监督的性能感知门控机制,以自适应地为定位效果较差的组保留证据引导,而当定位足够准确时则减少该引导,避免对细粒度边界细化的过度约束。在多个公开VTG基准上进行的大量实验验证了我们方法的有效性。
英文摘要
Large vision-language models (LVLMs) have achieved substantial performance gains in Video Temporal Grounding (VTG) through reinforcement learning (RL). However, existing methods primarily rely on outcome correctness rewards that evaluate only the final predicted intervals, leaving boundary-related visual evidence and its correspondence with timestamp predictions insufficiently constrained. In this paper, we delve into timestamp prediction and its underlying boundary-level visual evidence, showing prevalent misalignment between visual evidence and predicted timestamps across widely used benchmarks. To address this issue, we propose Competence-Aware Visual Boundary Evidence Alignment (CAVE), which augments localization optimization with boundary-specific visual evidence rewards to mitigate evidence-timestamp misalignment. Specifically, to explicitly represent the boundary-specific visual evidence, CAVE introduces boundary-specific evidence tokens and initializes their structured generation and distinct boundary semantics through a lightweight supervised warm-up. During RL, the visual boundary evidence alignment reward reinforces the visual attention of special evidence tokens within the ground-truth boundaries, thereby promoting alignment between visual evidence and temporal boundaries. Moreover, performance-aware gating for evidence supervision is designed to adaptively retain evidence guidance for poorly localized groups while reducing it once localization becomes sufficiently accurate to avoid over-constraining fine-grained boundary refinement. Extensive experiments on several public VTG benchmarks demonstrate the effectiveness of our method.