arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12203cs.CV

GeoFlow:基于几何对齐先验的高效驾驶视频生成

GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors

Jiazheng Liu, Hang Li, Jiawei Zhang, Jiahe Li, Xiaohan Yu, Shengyin Fan, Jin Zheng, Xiao Bai

首次发表
浏览论文内容

中文总结 AI 辅助

GeoFlow是一款新型驾驶视频生成框架,通过构建几何对齐先验分布替代标准高斯噪声初始化,可提升生成效率与少步生成质量,减少推理步骤。

中文摘要 AI 辅助

扩散模型、流匹配等生成模型在合成高保真驾驶视频方面展现出卓越能力,但因需要大量采样步骤而存在严重的推理延迟问题。我们认为这种低效性源于对标准高斯源分布的普遍依赖,该分布将连续帧初始化为独立高斯噪声,此范式忽略了驾驶视频固有的丰富时空相关性,迫使模型从噪声中重新生成前一帧中已有的确定性场景结构,这既存在计算冗余,又易引发几何不一致。为解决该问题,我们提出GeoFlow,一款利用显式几何先验实现高效驾驶视频生成的新型框架。我们不再从标准高斯噪声中采样,而是借助多视图几何和空间自适应噪声注入,构建几何对齐先验(GAP)分布作为起始点。这种初始化缩小了源分布与数据分布之间的差距,形成显著更短、更平滑的采样轨迹。大量实验表明,GeoFlow在训练和推理两方面均实现卓越效率:仅需对基线模型进行数小时微调,即可显著提升少步生成质量;完全收敛的训练则大幅减少了最先进视频生成所需的推理步骤数量。

英文摘要

Generative models like Diffusion Models and Flow Matching have demonstrated remarkable capabilities in synthesizing high-fidelity driving videos, but are severely constrained by high inference latency due to the requirement of extensive sampling steps. We argue that this inefficiency stems from the prevailing reliance on a standard Gaussian source distribution, where consecutive frames are initialized as independent Gaussian noise. This paradigm disregards the rich spatiotemporal correlations inherent in driving videos, compelling the model to regenerate deterministic scene structures existing in previous frames from noise, which is both computationally redundant and prone to geometric inconsistency. To address this problem, we propose GeoFlow, a novel framework designed to achieve efficient driving video generation by harnessing explicit geometric priors. Instead of sampling from standard Gaussian noise, we leverage multi-view geometry and spatially-adaptive noise injection to construct a Geometry-Aligned Prior (GAP) distribution as starting point. This initialization bridges the gap between source distribution and data distribution, yielding a significantly straighter and shorter sampling trajectory. Extensive experiments demonstrate that GeoFlow can achieve remarkable efficiency of both training and inference: merely several hours of fine-tuning on baseline models can significantly boost few-step generation quality, while fully converged training drastically reduces number of inference steps required for state-of-the-art video generation.

发表机构

  • Beihang University(北京航空航天大学)
  • Macquarie University(麦考瑞大学)
  • Tianyi Transportation Technology Co., Ltd(天宜交通科技有限公司)
  • Jiangxi Research Institute, Beihang University(北京航空航天大学江西研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑