arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13037cs.CV

基于推理时无梯度优化的空间接地文本到视频生成

Spatially-Grounded Text-to-Video Generation via Inference-Time Gradient-Free Optimization

发表机构ISIR - 索邦大学 · Obvious Research
查看机构详情
  • ISIR - Sorbonne Université(ISIR - 索邦大学)
  • Obvious Research

机构由 AI 辅助整理,请以论文原文为准。

Guillaume Jeanneret, Mathis Koroglu, Hugo Caselles-Dupré, Arnaud Dapogny, Matthieu Cord

首次发表
浏览论文内容

中文总结 AI 辅助

针对文本到视频生成的空间可控性挑战,提出无训练无梯度的GATO-Vid方法,通过解析求解交叉注意力分数实现精确空间引导,在提升定位精度的同时降低计算开销。

中文摘要 AI 辅助

扩散Transformer文本到视频模型已实现出色的合成质量,但细粒度空间可控性仍是重大挑战。现有无训练方法在空间接地生成(即将特定物体放置在指定位置)中能产生可靠整体结果,但依赖基于梯度的优化技术,会产生过高的计算开销,这一瓶颈在现代大规模架构中被放大。为解决此局限,我们提出Gradient-free Analytical Trajectory Optimization Video Generation(GATO-Vid),一种用于精确空间引导的新型无训练、无梯度方法。我们不依赖代价高昂的反向传播,而是引入替代交叉注意力分数并解析求解以获得精确闭式解;为使用该解析解,我们提出适配Transformer潜在空间拓扑流形的实时注入机制。实验表明,GATO-Vid在定位精度上显著优于现有基线,同时引入的计算开销极小。

英文摘要

Diffusion Transformer Text-to-Video models have achieved remarkable synthesis quality, yet fine-grained spatial controllability remains a significant challenge. While existing training-free methods produce solid overall results in spatially grounded generation, \ie, placing a specific object in a designated location, they rely on gradient-based optimization techniques that incur prohibitive computational overhead, a bottleneck amplified in modern large-scale architectures. To address this limitation, we present Gradient-free Analytical Trajectory Optimization Video Generation (GATO-Vid), a novel training-free and gradient-free approach for precise spatial guidance. Rather than relying on costly backward passes, we introduce an alternative cross-attention score and solve it analytically to obtain an exact, closed-form solution. To use our analytical solution, we propose an on-the-fly injection mechanism tailored to the topological manifold of the transformer's latent space. Our experiments demonstrate that GATO-Vid significantly outperforms existing baselines in localization accuracy while introducing minimal computational overhead.

补充信息

↑