发表机构
Shanghai Jiao Tong University; Zhejiang Normal University; Shandong Normal University; Tencent Hunyuan3D(上海交通大学; 浙江师范大学; 山东师范大学; 腾讯混元3D)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Flow3D-OPD两阶段后训练框架,通过多教师在线策略蒸馏解决3D几何生成中奖励定义和梯度干扰问题,实现全维度质量提升。
AI 中文摘要
近期基于流匹配扩散Transformer(DiT)构建的图像到3D生成模型能够产生高保真网格,但其后训练策略仍未得到充分探索。在强化学习中存在几个关键瓶颈:为3D几何质量定义全面奖励的固有困难,以及联合优化异构目标时出现的梯度干扰。受在线策略蒸馏(OPD)在大语言模型和图像生成中实用性的启发,我们提出了Flow3D-OPD,一个两阶段后训练框架,将多教师蒸馏引入3D几何生成。在第一阶段,我们利用半策略增强预训练模型的基础能力,然后设计一个智能体验证器用于3D几何质量评估。基于该验证器,我们可通过直接偏好优化(DPO)培养领域专精的教师模型。在第二阶段,我们通过带有硬任务路由采样和梯度累积的在线策略蒸馏,将异构专业知识整合到统一的学生模型中,这能缓解联合优化中的梯度干扰。无需依赖精细修改,我们简单而有效的设计在所有几何质量维度上取得了一致改进,并在平均指标上超越了所有教师模型。大量实验表明,我们的方法为3D生成中的强化学习提供了有效范式。
英文摘要
Recent image-to-3D generation models built on flow-matching diffusion Transformers (DiT) can produce high-fidelity meshes, yet their post-training strategy remains largely unexplored. There exist several critical bottlenecks in reinforcement learning: the inherent difficulty of defining comprehensive rewards for 3D geometric quality, and the gradient interference that arises when jointly optimizing heterogeneous objectives. Inspired by the practicability of on-policy distillation (OPD) in large language models and image generation, we propose \textbf{Flow3D-OPD}, a two-stage post-training framework that introduces multi-teacher distillation into 3D geometry generation. In the first stage, we utilize the semi-policy to enhance the foundational capability of the pretrained model and then design an agentic verifier for 3D geometric quality evaluation. Based on the verifier, we could cultivate domain-specialized teacher models via direct preference optimization (DPO). In the second stage, we consolidate heterogeneous expertise into a unified student model through on-policy distillation with hard task-routing sampling and gradient accumulation, which could mitigate the gradient interference in joint optimization. Without relying on elaborate modifications, our straightforward yet effective design achieves consistent improvements across all geometric quality dimensions and surpasses all teacher models in the average metric. Extensive experiments demonstrate that our approach provides an effective paradigm for reinforcement learning in 3D generation.