发表机构
Adobe Research(Adobe研究)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究逆渲染问题,提出完全黑箱框架FIDE,通过特征引导,利用视觉Transformer提取特征训练扩散模型,结合CMA进化策略优化,在多种逆问题上验证,显著提升收敛速度并逃离局部最小值。
AI 中文摘要
传统上,逆渲染通过可微渲染器和梯度下降来解决,这需要大量特定问题的工程设计,并且由于模糊性容易陷入局部最小值。无导数方法减轻了工程需求,但通常严重依赖于良好的问题初始化。在这项工作中,我们提出了特征信息扩散进化(FIDE),这是一个完全黑箱的框架,不需要梯度或特定初始化:渲染器被视为一个不透明的函数,唯一的要求是生成图像。我们的关键见解是特征引导:我们不是将每个候选渲染减少到一个标量损失值,而是使用视觉Transformer(ViT)从中提取密集的视觉特征。随后,我们使用这些特征来训练基于扩散的候选提案模型,使网络能够使用视觉线索来预测与目标图像匹配的参数。然后,通过CMA进化策略在闭环中对该扩散模型提出的候选解决方案进行优化,随着优化的进行不断缩小提案区域。我们在路径追踪、向量样条、Voronoi着色器和机器人等各种逆问题上进行了验证,并证明特征引导显著提高了收敛速度,超过了标量损失基线,并可靠地逃离了基于梯度的方法停滞的局部最小值。
英文摘要
Inverse rendering is traditionally solved via differentiable renderers and gradient descent, which requires substantial problem-specific engineering and is prone to getting stuck in local minima due to ambiguities. Derivative-free approaches alleviate engineering requirements, but often heavily depend on a good problem initialization. In this work, we propose Feature-Informed Diffusion Evolution (FIDE), a fully black-box framework that requires no gradients or specific initialization: the renderer is treated as an opaque function whose only requirement is to produce images. Our key insight is feature guiding: rather than reducing each candidate rendering to a scalar loss value, we use a Vision Transformer (ViT) to extract dense visual features from it. We subsequently use these features to train a diffusion-based candidate proposal model, allowing the network to use visual cues to predict parameters that would match the target image. The candidate solutions proposed by this diffusion model are then refined in a closed loop with a CMA evolution strategy, continuously narrowing the proposal region as optimization progresses. We validate across diverse inverse problems from path tracing, vector splines, Voronoi shaders, and robotics, and demonstrate that feature-guiding substantially improves convergence over scalar-loss baselines and reliably escapes local minima where gradient-based methods stall.