发表机构
Advanced Micro Devices, Inc.(超威半导体公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DAGS,通过解耦外观与几何的轻量条件引导冻结DiT,实现高保真、可控且时间稳定的生成式渲染,相比现有方法显著提升PSNR和时序稳定性。
AI 中文摘要
扩散变换器(DiTs)能从文本和图像条件生成高保真图像,但其输出具有较大方差,且对目标条件的忠实度高度依赖于条件的提供方式。我们提出DAGS,一种轻量级、无注意力、解耦的外观与几何条件方案,引导冻结的图像DiT生成高保真、高忠实度且可独立控制的渲染结果。两个小型卷积编码器每帧计算一次条件特征,并将其作为学习到的逐层逐元素残差注入图像令牌,避免了通过注意力堆叠条件的二次成本。由于控制和时序处理位于冻结主干之外,我们保留了其庞大的预训练先验,并消除了主干过拟合的风险。我们进一步添加了一个小型循环光照稳定器和一个免训练的时序引导项,与我们的条件方案相结合,将逐帧图像模型提升为流式渲染器。DAGS以路径追踪计算量的一小部分产生可控、高质量的渲染;它不是实时的,而是用计算换取可控性和质量。在匹配的1-spp+G-buffer输入上,逐帧DAGS相比实时降噪器Intel OIDN和扩散渲染器RGB<->X,PSNR分别提高+8.6 dB和+10.1 dB,同时在感知上(时序LPIPS闪烁)时间稳定性提高2.5-8倍。
英文摘要
Diffusion transformers (DiTs) generate high-fidelity images from text and image conditions, but their outputs carry large variance and their faithfulness to a desired target depends heavily on how the condition is supplied. We present DAGS, a lightweight, attention-free, disentangled appearance and geometry conditioning scheme that steers a frozen image DiT to produce high-fidelity, highly faithful, and independently controllable renders. Two small convolutional encoders compute conditioning features once per frame and inject them as a learned, per-layer, element-wise residual into the image tokens, avoiding the quadratic cost of stacking conditions through attention. Because control and temporal handling live outside the frozen backbone, we retain its vast pretrained prior and eliminate backbone-overfitting risk. We further add a small recurrent lighting stabilizer and a training-free temporal guidance term that, coupled with our conditioning, elevate a per-frame image model into a streaming renderer. DAGS produces controllable, high-quality renders at a fraction of the compute of path tracing; it is not real-time, trading compute for controllability and quality. On a matched 1-spp + G-buffer input, per-frame DAGS reconstructs +8.6 dB / +10.1 dB PSNR over the real-time denoiser Intel OIDN and the diffusion renderer RGB<->X while being 2.5-8x more temporally stable perceptually (temporal-LPIPS flicker).
Comments5 pages, 3 figures, 2 tables. Accepted to SIGGRAPH Asia 2026 Technical Communications