arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.22527cs.CV

轨迹强制:具有可控语义轨迹的结构优先生成

Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories

Merve Kocabas, Gege Gao, Bernhard Schölkopf, Andreas Geiger

首次发表
浏览论文内容

中文总结 AI 辅助

提出轨迹强制框架,将生成过程组织为从全局布局到细节的语义阶段,每个阶段产生可解码的潜在状态,支持局部编辑,实现结构优先的可控图像合成。

中文摘要 AI 辅助

扩散和基于流的生成模型能够生成高质量的图像,但其可控性主要局限于端点:用户指定条件并接收最终输出,而中间的生成动态过程是隐藏的。最近的方法开始利用生成顺序和过程分解来提高样本质量,但仍将中间状态视为内部计算而非可交互的对象。我们提出轨迹强制(TF),一种以轨迹为中心的框架,使生成路径显式、语义化且可编辑。TF将合成组织为一系列语义结构化的阶段,从全局布局逐步推进到对象、部件和细节级别的表示。每个阶段产生一个可解码的潜在状态,可以在下一阶段开始前进行检查、评估和局部编辑。为了实例化该路径,我们通过聚类预训练的视觉表示(如DINOv2)推导出从粗到细的教师层次结构,并在每个层次训练一个层次条件的一步流匹配模型。我们进一步引入了轨迹感知的指标,用于衡量结构一致性和局部可控性,超越了端点质量指标如FID。实验表明,TF在保持竞争性样本质量的同时,暴露了连贯的中间状态,并支持跨语义层次的局部编辑。通过将焦点从最终图像转移到生成路径本身,TF开辟了一条通向可控、轨迹感知的图像合成之路。

英文摘要

Diffusion and flow-based generative models produce strong images, yet their controllability remains largely endpoint-centric: users specify conditions and receive final outputs, while the intermediate generative dynamics remain hidden. Recent methods have begun to exploit generation order and process decomposition to improve sample quality, but still treat intermediate states as internal computation rather than objects for interaction. We propose Trajectory Forcing (TF), a trajectory-centric framework that makes the generation path explicit, semantic, and editable. TF organizes synthesis as a sequence of semantically structured stages, progressing from global layout to object-, part-, and detail-level representations. Each stage produces a decodable latent state that can be inspected, evaluated, and locally edited before the next stage begins. To instantiate this path, we derive coarse-to-fine teacher hierarchies by clustering pretrained visual representations such as DINOv2, and train a hierarchy-conditioned one-step flow-matching model at each level. We further introduce trajectory-aware metrics that measure structural consistency and local controllability beyond endpoint quality metrics such as FID. Experiments show that TF achieves competitive sample quality while exposing coherent intermediate states and supporting localized edits across semantic levels. By shifting the focus from final images to the generative path itself, TF opens a route toward controllable, trajectory-aware image synthesis.

发表机构

  • University of Tübingen(图宾根大学)
  • ETH Zürich(苏黎世联邦理工学院)
  • Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
  • ELLIS Institute(ELLIS研究所)
  • Tübingen AI Center(图宾根人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑