基于生成先验蒸馏的无配对监督3D编辑学习
Learning 3D Editing without Paired Supervision via Generative Prior Distillation
- Beihang University(北京航空航天大学)
- Central University of Finance and Economics(中央财经大学)
- VAST
- Beijing Key Laboratory of Intelligent Creative Content Generation and Immersive Experience(智能创意内容生成与沉浸式体验北京市重点实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出基于生成先验蒸馏的无配对监督3D编辑框架,通过多模态先验与3D感知正则化实现高质量3D编辑,性能优于现有方法。
AI中文摘要:
指令引导的3D编辑是交互式内容创作的核心需求,但面临着高质量配对训练数据严重匮乏的瓶颈。现有方法试图绕过这一问题,要么依赖耗时的测试时优化,要么基于复杂流程构建的伪配对进行训练,这往往会引入结构漂移和几何伪影。本文提出一种新型框架,通过生成先验蒸馏学习无配对3D监督的前馈3D编辑模型。核心思路不依赖真实3D配对,而是将强大基础模型的视觉、语义和几何知识直接蒸馏到3D编辑模型中。具体而言,通过可微渲染管线,我们用两种互补信号监督3D表示:主编辑视图下图像编辑模型提供的2D视觉先验,以及新视图下视觉语言模型提供的语义先验,以确保严格遵循指令并保留源身份。关键在于,为解决2D投影监督固有的几何崩溃和多视图不一致问题,我们引入了3D感知分布匹配正则化项,作为几何先验,该术语在3D隐空间中运行,约束编辑输出保持在预训练图像转3D教师模型定义的真实3D资产流形内。大量实验表明,我们的方法实现了卓越的指令保真度和跨视图一致性,显著优于最先进的基线方法。该项目可在this https URL获取。
英文摘要:
Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottleneck: the severe scarcity of high-quality paired training data. Existing approaches attempt to bypass this by either relying on slow test-time optimization or training on pseudo-pairs constructed via complex pipelines, which often introduce structural drift and geometric artifacts. In this paper, we propose a novel framework that learns feed-forward 3D editing without paired 3D supervision via Generative Prior Distillation. Instead of relying on ground-truth 3D pairs, our core idea is to distill visual, semantic, and geometric knowledge from powerful foundation models directly into a 3D editing model. Specifically, through a differentiable rendering pipeline, we supervise the 3D representation using two complementary signals: a 2D visual prior from an image editing model at the main editing view, and a semantic prior from a Vision-Language Model at novel views to ensure strict instruction following and source identity preservation. Crucially, to address the geometric collapse and multi-view inconsistencies inherent in 2D projection supervision, we introduce a 3D-aware Distribution Matching regularization. Acting as a geometric prior, this term operates in the 3D latent space, constraining the edited output to remain within the manifold of realistic 3D assets defined by a pretrained image to 3D teacher model. Extensive experiments demonstrate that our method achieves superior instruction fidelity and cross-view consistency, significantly outperforming state-of-the-art baselines. Our project is available at: https://github.com/thiamine128/PriorEdit3D.