arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.08016cs.CVcs.GR

LightCrafter:用于可控且一致的重光照的PBR条件视频扩散细化

LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting

  • Carnegie Mellon University(卡内基梅隆大学)
  • University of Toronto(多伦多大学)
  • Bosch Research(博世研究)

机构由 AI 辅助整理,请以论文原文为准。

Zixin Guo, Yehonathan Litman, Yifeng He, John Miller, Chuhan Chen, Deva Ramanan

AI总结:

研究视频重光照问题,提出LightCrafter混合管道,将其重表述为代理视频转换,利用PBR渲染并结合光度先验,在真实世界重光照基准上优于现有技术,还贡献合成基准及相关资源。

AI中文摘要:

视频重光照需要在长时的时间一致性与对光传输的物理基础理解之间取得平衡,这依赖于对材料、几何形状和光照等内在场景属性的准确估计。现有方法有两种范式:一是通过逆渲染重建视频的光度属性并通过正向渲染将其重光照到目标光照下,但存在重建噪声且难以处理全局光照等难以建模的效果;二是将任务作为基于重光照目标的生成式视频到视频转换,但存在重光照控制和时间稳定性受限等问题。我们提出LightCrafter,一种混合管道,将视频重光照重新表述为代理视频的视频转换,将输入在目标光照下的PBR渲染转换为最终目标,能实现更精细的光照控制并自然提供长时时间一致性。单独的PBR渲染已优于一些现有技术,但仍难以处理全局光照等效果,我们通过在合成视频对和真实世界未配对视频上对CogVideoX进行后训练来利用视频生成模型中的光度先验。我们在现有真实世界重光照基准上优于先前的最先进技术,并贡献了一个合成基准用于进一步分析。我们将发布数据集、基准、指标和代码。

英文摘要:

Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurate estimation of intrinsic scene properties such as materials, geometry, and illumination. Existing methods follow two paradigms: (1) reconstruct a video's photometric properties via inverse rendering and relight them to a target illumination via forward rendering, using physically-based rendering (PBR) or a neural renderer; these suffer from noisy reconstructions and struggle with hard-to-model effects such as global illumination. (2) Frame the task as generative video-to-video translation conditioned on relighting targets (a target environment map or text); this limits relighting control and temporal stability, since diffusion models struggle to translate long-form videos, and is constrained by the availability of input/relit training pairs. We propose LightCrafter, a hybrid pipeline that reformulates video relighting as video translation of a proxy video: rather than translating the input video directly to the target, we translate a PBR rendering of the input under the target illumination to the final target. This bakes illumination targets into the PBR proxy, removing the need to teach the diffusion model illumination concepts like environment maps, and enables more intricate lighting control while naturally providing long-form temporal consistency. We show PBR renders alone already outperform some prior art but struggle with effects like global illumination; to capture these, we leverage photometric priors in video generation models by post-training CogVideoX on synthetic video pairs and real-world unpaired videos. We outperform prior state-of-the-art on existing real-world relighting benchmarks and contribute a synthetic benchmark for further analysis. We will release our dataset, benchmark, metrics, and code.

补充信息

↑