发表机构
Tencent Hunyuan; Shanghai Jiao Tong University; Communication University of China(腾讯混元; 上海交通大学; 中国传媒大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对三维生成中现有强化学习方法效果有限的问题,提出前向过程强化方法动态归位优化(DHO),通过最小成本吸引匹配和时间感知动态校正引导负轨迹朝向正样本,并构建Flow3D-Pro框架,实验证明其优于现有方法。
AI 中文摘要
流匹配是三维生成的核心,然而在实践中,其强化学习(RL)方法大多是从二维视觉生成中改编而来。代表性的DPO、GRPO和NFT风格目标函数,当应用于负轨迹时,主要引导预测速度远离相应方向,而没有明确指定一个朝向优选样本的目标速度场。在三维生成中,受预训练模型能力、展开多样性和奖励分布复杂性的限制,直接应用这些RL方法在几何质量上的提升有限。我们提出了一种前向过程强化学习方法——动态归位优化(DHO),它将负轨迹优化重新表述为正样本吸引引导的动态归位。具体而言,最小成本吸引匹配(MAM)为每个负样本分配一个不同的正目标,而时间感知动态校正(TDC)则利用剩余时间感知的校正速度将其轨迹重定向到目标。基于异步在线DHO,我们开发了Flow3D-Pro,一个图像到三维几何生成框架。实验表明,在三维生成中,DHO优于代表性的DPO、GRPO和NFT风格目标函数,而Flow3D-Pro生成的几何质量高于现有网格生成方法。
英文摘要
Flow matching is central to 3D generation, yet in practice its reinforcement learning (RL) methods are largely adapted from 2D visual generation. Representative DPO-, GRPO-, and NFT-style objectives, when applied to negative trajectories, mainly steer predicted velocities away from the corresponding directions without explicitly specifying a target velocity field toward preferred samples. In 3D generation, constrained by pretrained model capabilities, rollout diversity, and reward-distribution complexity, directly applying these RL methods yields limited gains in geometric quality. We introduce a forward-process RL method \textbf{Dynamic Homing Optimization (DHO)}, which reformulates negative-trajectory optimization as positive-sample attraction-guided dynamic homing. Specifically, Minimum-Cost Attractive Matching (MAM) assigns each negative sample a distinct positive target, and Time-Aware Dynamic Correction (TDC) then redirects its trajectory toward the target using a remaining-time-aware corrective velocity. Building on asynchronous online DHO, we develop \textbf{Flow3D-Pro}, an image-to-3D geometry generation framework. Experiments show that DHO outperforms representative DPO-, GRPO-, and NFT-style objectives in 3D generation, while Flow3D-Pro produces higher-quality 3D geometry than existing mesh generation methods.