arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28549cs.CVcs.AI

作为几何学习者的视频生成模型

Video Generative Models as Geometry Learner

Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出GeoNeXt方法,将预训练视频生成模型用作统一数据高效的几何估计框架,以零样本方式在少数据下实现单目深度和表面法向量估计,性能优于同类生成式方法,可媲美数据量超100倍的判别式SOTA方法。

中文摘要 AI 辅助

近期用于几何估计的生成方法会适配预训练的图像扩散模型,并将该任务视为图像条件生成任务。这类方法利用现成的图像扩散模型,要么(i)独立训练特定任务的几何模型(用于深度和表面法向量估计),错失了探索这些几何目标内在关联的机会;要么(ii)联合微调修改后的图像扩散骨干网络(例如修改自注意力机制),这通常需要大量标注数据。为以合理方式克服这些局限,我们将预训练的视频生成模型重新用作几何估计的统一且数据高效的框架,创新性地将其表述为下一帧预测任务。我们的方法GeoNeXt自然继承了视频模型的结构化知识和更丰富的先验,同时进一步将其适配为图像与几何目标(图像↔几何)的联合建模,实现了更数据高效且有效的几何学习。大量实验验证了我们的方法在不同数据集上的零样本单目深度和表面法向量估计性能,其表现优于此前的特定任务及统一生成式竞争对手,同时使用的训练数据量显著更少。值得注意的是,我们的方法可与训练数据量超100倍的判别式SOTA方法相媲美,甚至在多个基准上表现更优。

英文摘要

Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry models (for depth and surface normal estimation) independently, losing the opportunity of exploring the intrinsic correlation of these geometric targets, or (ii) jointly fine-tune modified image diffusion backbones (e.g., altered self-attention), which typically demands substantial labeled data. To overcome these limitations in a principled fashion, we repurpose pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task. Our method, GeoNeXt, inherits naturally structured knowledge and richer priors from the video model, while further adapting them for joint modeling of images and geometry targets (image <-> geometry), enabling more data efficient and effective learning of geometry. Extensive experiments validate our method for zero-shot monocular depth and surface normal estimation across diverse datasets, outperforming both previous task-specific and unified generative competitors while using substantially less training data. Notably, our method rivals discriminative state-of-the-art approaches trained on over 100x more data and even standouts on several benchmarks.

发表机构

  • University of Surrey(萨里大学)
  • Imperial College London(伦敦帝国学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑