AI 中文总结
研究模拟到真实翻译中弥合合成与真实域外观差距并保持结构语义一致的问题,提出小波相位扩散方法,通过双树复小波包变换域操作及低频随机化,在未配对数据训练,提升了图像和视频翻译的真实感与一致性。
AI 中文摘要
模拟到真实的翻译必须弥合合成域和真实域之间的外观差距,同时保持结构和语义一致性。基于条件的方法实现了空间对齐,但引入了计算成本高昂的控制模块。配对数据方法实现了真实感,但依赖复杂的合成管道,常常改变场景几何和语义。无需训练的编辑方法避免了这两个限制,但缺乏学习到的外观先验,限制了感知质量。最近提出的相位保留扩散是一个有前途的替代方案,但傅里叶域公式受全局频谱耦合的限制。这种耦合会导致空间伪影,如振铃和边界泄漏,从而降低结构和语义一致性。我们引入了小波相位扩散,通过两个组件来解决这个问题。首先,我们在双树复小波包变换域中操作,其局部小波包能够进行空间自适应相位注入而无全局频谱干扰。其次,低频随机化(LFR)取代低频包,使模型与合成光照先验解耦,并实现分布内的真实世界外观。这两个组件都在未配对的开放域数据上进行训练,并且引入可忽略不计的推理开销。空间局部性还实现了实例级翻译,其中单个对象或区域被独立地翻译成逼真的外观,而周围场景保持不变。在vKITTI→KITTI图像翻译中,我们的方法在真实感和语义一致性方面优于先前方法,同时保持有竞争力的结构对齐。对于CARLA视频翻译,我们的方法接近配对数据方法的真实感,同时分别将VLM规划器的ADE和FDE降低了5.4%和5.1%。
英文摘要
Simulation-to-reality translation must bridge the appearance gap between synthetic and real domains while preserving structural and semantic consistency. Conditioning-based methods achieve spatial alignment but introduce computationally expensive control modules. Paired-data methods achieve realism but rely on complex synthesis pipelines, often altering scene geometry and semantics. Training-free editing methods avoid both constraints but lack a learned appearance prior, limiting their perceptual quality. Recently proposed phase-preserving diffusion presents a promising alternative, but Fourier-domain formulations are constrained by global spectral coupling. This coupling induces spatial artifacts such as ringing and boundary leakage, thereby degrading structural and semantic consistency. We introduce Wavelet Phase Diffusion, which addresses this through two components. First, we operate in the Dual-Tree Complex Wavelet Packet Transform domain, whose localized wavelet packets enable spatially adaptive phase injection without global spectral interference. Second, Low-Frequency Randomization (LFR) replaces the low-frequency packet, decoupling the model from the synthetic illumination prior and enabling in-distribution real-world appearance. Both components train on unpaired open-domain data, and introduce negligible inference overhead. The spatial locality further enables instance-level translation, where individual objects or regions are translated to photorealistic appearance independently while the surrounding scene remains untranslated. On vKITTI $\to$ KITTI image translation, ours outperforms prior methods in realism and semantic consistency while maintaining competitive structural alignment. For CARLA video translation, ours approaches the realism of paired-data methods while reducing VLM planner ADE and FDE by $5.4\%$ and $5.1\%$, respectively.
CommentsProject Page: https://kit-mrt.github.io/Wavelet-Phase-Diffusion/