arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07414cs.CVcs.AIcs.GRcs.LGcs.MM

RelightFormer:用于多视图物体重照明的前馈生成式Transformer

RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting

Hejun Wang, Jinxi Li, Junwei Jiang, Shiwei Mao, Hu Cheng, Shouwang Huang, Bo Yang

首次发表
浏览论文内容

中文总结 AI 辅助

提出前馈生成式Transformer RelightFormer,通过潜在照明模块和置换不变编码实现单/多视图物体重照明,并构建大规模LOD数据集,达到最先进效果与零样本泛化。

中文摘要 AI 辅助

图像重照明传统上通过复杂的逆渲染管线处理,这些管线受困于不适定优化问题,或依赖忽略理解3D几何和材质交互所必需的关键多视图线索的单图像生成模型。为解决这些局限,我们提出一种前馈生成式Transformer,用于直接进行单视图和多视图图像重照明,完全绕过了显式的固有属性估计。该架构改编自视频基础模型,包含一个潜在照明模块,通过交叉注意力将目标环境图动态注入空间特征。此外,我们采用置换不变的位置编码,以对称方式处理无序的多视图输入,避免顺序偏差。为训练这一鲁棒的数据驱动模型,我们构建了大规模Laval Objaverse数据集(LOD),包含9万个物体和3.9万种独特照明。大量实验表明,在单视图、多视图和新视图重照明任务中,我们的方法实现了最先进的视觉质量、逼真的重照明效果以及强大的零样本泛化能力。

英文摘要

Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.

发表机构

  • Shenzhen Research Institute, The Hong Kong Polytechnic University(香港理工大学深圳研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑