复制所见,生成未见:无训练异常感知视频修复
Copy What Is Seen, Generate What Is Not: Training-Free Anomaly-Aware Video Restoration
浏览论文内容
中文总结 AI 辅助
提出AVR,用冻结预训练模型弥合无训练异常检测与视频修复的差距,仅在无可复制证据处生成内容,在监控数据集上实现高保真修复并优于先检测后生成流程。
中文摘要 AI 辅助
一个检测异常的监控系统通常也需要修复视频片段,然而这两个任务却被孤立研究:无训练的异常检测器止步于输出分数或标签,而无训练的视频编辑则响应用户提示而非检测器。本文提出AVR(异常感知视频修复),仅使用冻结的预训练模型弥合这一差距,并且只在片段无法提供可复制证据的位置生成内容。首先,运动证据将开放词汇提议门控为时空掩码。然后,从片段计算得到的背景先验填充异常所暴露的每个像素,让扩散模型仅合成任何帧都未显示的内容,最后,一个冻结的验证器逐片段决定信任经典修复器、先验锚定修复器还是背景条件修复器。在三个监控数据集上,分别在全参考异常注入和真实异常条件下进行的大量实验表明,AVR在掩码已知时实现全帧保真度,在编辑区域内匹敌三个训练过的视频修复器,并在其自身生成的掩码上优于先检测后生成的流程,同时抑制残余异常和自由扩散的闪烁。
英文摘要
A surveillance system that detects an anomaly often has to repair the footage as well, yet the two tasks are studied in isolation: training-free anomaly detectors stop at a score or a label, while training-free video editing answers to a user prompt rather than to a detector. This paper proposes AVR (Anomaly-aware Video Restoration), which closes that gap with frozen pretrained models alone and generates content only where the clip offers no evidence to copy. Motion evidence first gates open-vocabulary proposals into spatio-temporal masks. A background prior computed from the clip then fills every pixel the anomaly ever uncovers, leaving diffusion to synthesize only what no frame showed, and a frozen verifier decides per clip whether to trust a classical, a prior-anchored, or a background-conditioned restorer. Extensive experiments on three surveillance datasets, under both full-reference anomaly injection and real anomalies, show that AVR leads full-frame fidelity under oracle masks, matches three trained video inpainters inside the edited region, and outperforms a detect-then-generate pipeline on the masks it produces itself, while suppressing both the residual anomaly and the flicker of free diffusion.
发表机构
- New York University(纽约大学)
- University of Technology Sydney(悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。