W2Rep:通过观察世界变化学习视觉表征
W2Rep: Learning Visual Representations by Watching the World Change
浏览论文内容
中文总结 AI 辅助
W2Rep通过掩码特征预测框架,利用视频中的场景变化监督视觉编码器,在不牺牲视频表征能力的同时提升单图像特征,实验证明其在冻结和微调识别任务中均优于基线。
中文摘要 AI 辅助
图像捕捉世界某一时刻的状态,而视频则揭示其变化过程。图像自监督学习从单一时刻提取空间结构,而视频方法通常在一个由多帧联合计算得到的表征内部学习时间关系。我们探究能否通过观察场景变化来改进单张图像可用的特征,同时不牺牲表征视频的能力。为此,我们提出W2Rep,一种掩码特征预测框架,其中独立编码的源图像参与同一时刻或另一时刻的预测。预测器以可见视频上下文、查询位置以及源图像与目标之间的有符号时间间隔为条件。这使得跨帧目标具有两个互补作用:图像路径学习随时间保持有用的特征,而视频路径必须收集源图像中缺失的证据。在多种模型规模和下游任务中,W2Rep在我们的比较协议下提升了冻结和微调识别性能,而联合视频编码相比逐帧聚合带来了进一步增益。控制实验表明,这些增益依赖于直接更新源图像特征以及同时使用视频上下文和时间位移。总体而言,视频中的变化可以监督一个视觉编码器,其表征在图像或视频粒度下均保持有用。代码可在~\href{this https URL}{this https URL}获取。
英文摘要
Images capture the world at one moment, whereas video reveals how it changes. Image self-supervision learns spatial structure from a single moment, while video methods commonly learn temporal relationships inside a representation computed jointly from several frames. We ask whether watching a scene change can instead improve features available from one image without sacrificing the ability to represent video. We introduce W2Rep, a masked feature-prediction framework in which an independently encoded source image participates in prediction at the same or another moment. The predictor is conditioned on visible video context, the queried location, and the signed time interval between source and target. This gives the cross-frame objective two complementary roles: the image path learns features that remain useful across time, while the video path must gather evidence that is missing from the source image. Across model scales and downstream tasks, W2Rep improves frozen and fine-tuned recognition under our comparison protocol, while joint video encoding provides further gains over frame-wise aggregation. Controlled experiments show that these gains depend on directly updating the source-image features and on using both video context and temporal displacement. Overall, change across a video can supervise a visual encoder whose representations remain useful at either image or video granularity. Code is available at~\href{https://wenooi.github.io/W2Rep}{https://wenooi.github.io/W2Rep}.
发表机构
- Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
- Nankai University(南开大学)
- Shenzhen University(深圳大学)
机构由 AI 辅助整理,请以论文原文为准。