FixAnything:基于视频生成先验的3D一致性渲染优化
FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors
浏览论文内容
中文总结 AI 辅助
FixAnything是一款复用预训练视频生成模型、仅需轻量微调的单一模型,可修复3DGS、NeRF等四种3D表示的渲染伪影,实现3D一致性渲染,替代多个专用优化管线。
中文摘要 AI 辅助
使用3D场景表示(如高斯溅射(3DGS)、神经辐射场(NeRF)、网格甚至点云)渲染视图时,若输入视图稀疏或目标视图与输入视图距离较远,会产生伪影。近期研究采用基于扩散的生成先验缓解此类伪影,但这类方法针对特定场景表示,需定制架构或大量重新训练。本文提出FixAnything,一款用于修复各类渲染伪影的单一模型,通过复用预训练视频生成模型实现,仅做 minimal 修改与轻量微调即可利用其隐式多视图先验。核心思路是:即使渲染存在噪声,序列仍保留相机运动与粗略场景结构,因此可将清理任务表述为视频到视频的转换;为控制需保留的场景结构,引入表示干净像素的二值掩码,使模型锚定高质量输入(如训练视图)的同时优化其余部分;为促使FixAnything生成支持下游重建的3D一致性渲染,采用通过运动恢复结构(SfM)得到的相机位姿精度作为直接偏好优化(DPO)的奖励信号。在四种不同3D表示上的实验表明,FixAnything通过轻量微调持续提升渲染质量,证明单一通用视频先验可替代多个专用优化管线,该框架的简洁性使其无需架构重新设计即可立即采用未来更强的视频模型。
英文摘要
Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces artifacts when input views are sparse or target views lie far from the input. Recent work mitigates these artifacts using diffusion-based generative priors, but is specialized to individual representations and require custom architectures or extensive retraining. We present FixAnything, a single model for fixing a wide range of rendering artifacts. It does so by repurposing a pretrained video generative model, leveraging its implicit multi-view priors with only minimal modification and lightweight finetuning. Our key insight is that even noisily-rendered sequences preserve camera motion and coarse scene structure, allowing cleanup to be formulated as video-to-video translation. To control what scene structure should be preserved, we introduce a binary mask denoting the clean pixels, enabling the model to anchor its output to high-quality inputs (e.g. training views) while refining the rest. To encourage FixAnything to produce 3D-consistent renderings that support downstream reconstruction, we use camera pose accuracy (recovered via structure-from-motion) as a reward signal for direct preference optimization (DPO). Across four distinct 3D representations, FixAnything consistently improves rendering quality with lightweight finetuning, demonstrating that a single generalist video prior can replace multiple specialist refinement pipelines. The simplicity of the framework enables immediate adoption of stronger future video models without architectural redesign.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。