任意视频时间定位器的共形覆盖保证
Conformal Coverage Guarantees for Any Video Temporal Grounder
浏览论文内容
中文总结 AI 辅助
针对视频时间定位器的模糊性问题,提出与模型无关的COVER包装器,可保证输出区域以至少1-α的概率包含真实时刻,在多基准和定位器上验证了其覆盖度符合目标值。
中文摘要 AI 辅助
连续视频中的事件边界具有模糊性:对相同的查询-视频对进行重新标注,独立标注者在很大比例的样本上标记的时刻重叠度不足一半。因此,视频时间定位(video temporal grounding)的真实值是区间上的分布,但每个定位器(grounder)仅返回单个区间,且不提供任何可靠性说明,导致在部署时,错误区间与正确区间无法区分。COVER 改变了输出对象:它是一种事后的、与模型无关的包装器,可将任意定位器(无论是经过训练的定位器还是黑盒视频-语言模型)转换为能输出时间区域的模型,该区域包含真实时刻的概率至少为 $1-\alpha$,其实现方式是在保留的标签上校准时间非一致性得分的分位数,并以此量扩大基础预测区间。该保证在可交换性下是有限样本且与分布无关的,既不需要重新训练,也不需要白盒访问。我们提供两类得分:针对输出区间的定位器的双侧边界扩充分数,以及针对输出相关性信号的定位器的超水平集得分;并开发了针对定位任务的理论,该理论可界定认证区域的大小、覆盖度在以事件长度为条件时的保持情况,以及当单个视频中的时刻破坏可交换性时覆盖度的下降情况。在三个基准和五个定位器上,实际覆盖度与目标值相符,且校准过程揭示了点指标所隐藏的信息。
英文摘要
Event boundaries in continuous video are ambiguous: re-annotate the same query-video pair and independent annotators mark moments that overlap by less than half on a large fraction of samples. The ground truth for video temporal grounding is therefore a distribution over intervals, yet every grounder returns a single interval with no statement of reliability, so at deployment a wrong interval is indistinguishable from a right one. COVER changes the output object: a post-hoc, model-agnostic wrapper that turns any grounder, a trained localizer or a black-box video--language model, into one that emits a temporal region containing the true moment with probability at least $1-α$, by calibrating the quantile of a temporal nonconformity score on held-out labels and widening the base prediction by that amount. The guarantee is finite-sample and distribution-free under exchangeability, and requires neither retraining nor white-box access. We give two score families, a two-sided boundary-widening score for grounders that emit an interval and a super-level-set score for grounders that emit a relevance signal, and develop theory specific to grounding that bounds how large the certified region becomes, when coverage survives conditioning on event length, and how it degrades when moments from one video break exchangeability. Across three benchmarks and five grounders, realized coverage tracks the target, and calibration exposes what point metrics hide.