发表机构
Pukyong National University(釜庆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出扩散轨迹建模(DTM),将扩散过程中每个空间补丁的时间演化视为轨迹,用于语义对应,在SPair-71k等基准上取得强性能,开辟了利用扩散表示的新视角。
AI 中文摘要
扩散模型通过迭代扩散过程生成图像,近期研究表明,该过程中产生的中间特征图包含丰富的视觉表示,因此被广泛应用于各种下游任务。然而,大多数现有方法仅限于使用特定时间步的单一特征图,或跨多个时间步聚合特征图。我们观察到,扩散过程中的中间表示沿时间轴形成有意义的轨迹。特别是,每个空间补丁的表示在整个生成过程中逐步演化,编码了仅从静态快照难以捕捉的语义。这一观察促使我们需要将扩散表示视为时间结构化的轨迹,而非静态快照。为此,我们提出了扩散轨迹建模(DTM),一个将每个空间补丁的时间演化解释为轨迹并利用其进行语义对应的框架。通过有效建模跨多个时间步生成的补丁级轨迹,DTM捕捉了先前方法未设计捕捉的对应线索。我们进一步通过实验证明,空间对应的补丁在整个扩散过程中形成相似的轨迹模式,表明扩散的时间轴携带语义信息。在SPair-71k、SPair-U和AP-10K上的实验表明,DTM取得了强劲的性能,为从轨迹中心视角利用扩散表示提供了新视角。
英文摘要
Diffusion models generate images through an iterative diffusion process, and recent studies have demonstrated that the intermediate feature maps produced during this process contain rich visual representations, leading to their adoption across a variety of downstream tasks. However, most existing approaches are limited to either using a single feature map at a specific timestep or aggregating feature maps across multiple timesteps. We observe that intermediate representations in the diffusion process form meaningful trajectories along the time axis. In particular, the representation of each spatial patch evolves progressively throughout the generative process, encoding semantics that are difficult to capture from static snapshots alone. This observation motivates the need to treat diffusion representations as temporally structured trajectories rather than static snapshots. To this end, we propose Diffusion Trajectory Modeling (DTM), a framework that interprets the temporal evolution of each spatial patch as a trajectory and leverages it for semantic correspondence. By effectively modeling patch-wise trajectories generated across multiple timesteps, DTM captures correspondence cues that prior methods are not designed to capture. We further demonstrate empirically that spatially corresponding patches form similar trajectory patterns throughout the diffusion process, suggesting that the temporal axis of diffusion carries semantic information. Experiments on SPair-71k, SPair-U and AP-10K show that DTM achieves strong performance, presenting a new perspective for exploiting diffusion representations from a trajectory-centric viewpoint.
CommentsAccepted to BMVC2026