arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ScaleVid:基于无网格推理的几何感知视频目标缩放

ScaleVid: Geometry-Aware Video Object Scaling with Mesh-Free Inference

Youze Huang, Penghui Ruan, Bojia Zi, Xianbiao Qi, Shihao Zhao, Rong Xiao

arXiv 2608.12232首次发表:更新:

AI 中文总结

ScaleVid提出无网格推理的两阶段训练框架,实现几何感知视频目标缩放,在野外视频评估中,其几何一致性、前景保真度等优于需显式3D重建的方法,推理更实用。

AI 中文摘要

几何感知视频目标缩放旨在沿以目标为中心的轴对目标进行各向异性缩放,同时保持几何合理性、时间一致性和背景一致性。现有文本引导方法主要在2D图像平面上操作,而深度引导方法提供的控制较为粗略,基于网格的方法则需要昂贵的3D重建。我们提出了一种渐进式两阶段训练框架,该框架将几何感知前景变换与背景保留及逼真视频合成解耦,推理时无需网格-像素对齐和显式3D重建。在两个阶段中,均从真实视频构建几何扰动的伪源,同时保留原始完整视频作为重建目标。第一阶段使用平面变换学习鲁棒的前景-背景合成,第二阶段引入以目标为中心的3D变形引导以实现几何感知缩放。这种伪源重建公式可实现无配对真实世界缩放目标的真实视频合成。我们构建了互补的配对几何和真实背景基准,并对野外视频进行评估。大量实验表明,与需要显式3D重建的方法相比,本方法具有更优的几何一致性、前景保真度和背景保留能力,同时推理速度更快、更实用。

英文摘要

Geometry-aware video object scaling aims to anisotropically resize the object along object-centric axes while preserving geometric plausibility, temporal coherence, and background consistency. Existing text-guided methods mainly operate in the 2D image plane, while depth-guided approaches provide coarse control and mesh-based methods require costly 3D reconstruction. We present a progressive two-stage training framework that decouples geometry-aware foreground transformation from background preservation and realistic video composition, without mesh-pixel alignment and explicit 3D reconstruction at inference. In both stages, geometrically perturbed pseudo-sources are constructed from real videos, while the original complete videos are retained as reconstruction targets. The first stage uses planar transformations to learn robust foreground-background composition, whereas the second introduces object-centric 3D deformation guidance for geometry-aware scaling. This pseudo-source reconstruction formulation enables real-video synthesis without paired real-world scaling targets. We construct complementary paired-geometry and real-background benchmarks and further evaluate on in-the-wild videos. Extensive experiments demonstrate superior geometric consistency, foreground fidelity, and background preservation, together with faster and more practical inference than methods requiring explicit 3D reconstruction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑