arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03423cs.CV

OuroReward:文本到3D生成中强化学习的顺序奖励调度

OuroReward: Sequential Reward Scheduling for Reinforcement Learning in Text-to-3D Generation

Bingyang Cui, Yujie Zhang, Yiling Xu, Yunfeng Guan

首次发表
浏览论文内容

中文总结 AI 辅助

提出OuroReward,一种干扰感知的顺序奖励调度策略,通过构建循环优化路径并自适应推进,结合AdaSelect提示选择,提升文本到3D生成中强化学习的多维度质量。

中文摘要 AI 辅助

文本到3D(T23D)生成的强化学习(RL)需要在多个质量维度上进行优化,例如语义对齐和纹理清晰度。现有方法通常通过多个奖励的聚合同时优化这些维度,而没有显式建模维度间的依赖关系。这可能导致优化不平衡以及冲突维度间的持续干扰。为解决这一局限,我们提出OuroReward,一种用于T23D RL的干扰感知的顺序奖励调度策略。OuroReward首先估计维度间的成对依赖关系,并构建一个最小化累积干扰的循环优化路径。通过纳入尾到头依赖,该循环捕获了整个调度中的全局兼容性。然后,OuroReward将循环转换为单遍序列,并从具有最低总干扰的维度开始优化。训练并非为每个维度奖励分配固定的优化预算,而是根据当前奖励的剩余优化空间自适应地决定何时推进到下一个奖励。我们进一步引入AdaSelect,一种自适应提示选择策略,用于识别与模型当前能力相符的可靠且信息丰富的提示。通过将策略更新聚焦于这些提示,AdaSelect有效提升了训练稳定性。在不同T23D模型和RL算法上的大量实验表明,我们的框架在多个维度上持续提高了生成质量。

英文摘要

Reinforcement learning (RL) for Text-to-3D (T23D) generation requires optimization across multiple quality dimensions such as semantic alignment and texture clarity. Existing methods typically optimize these dimensions simultaneously through multiple reward aggregation, without explicitly modeling inter-dimension dependencies. This can cause imbalanced optimization and persistent interference among conflicting dimensions. To address this limitation, we propose OuroReward, an interference-aware sequential reward scheduling strategy for T23D RL. OuroReward first estimates pairwise dependencies among dimensions and constructs a cyclic optimization path that minimizes cumulative interference. By incorporating the tail-to-head dependency, the cycle captures global compatibility across the entire schedule. Then, OuroReward converts the cycle into a one-pass sequence, and starts optimization from the dimension with the lowest aggregate interference. Rather than assigning a fixed optimization budget to each dimension-wise reward, training adaptively determines when to advance to the next reward according to the remaining optimization headroom of the current one. We further introduce AdaSelect, an adaptive prompt selection strategy that identifies reliable and informative prompts aligned with the model's current capability. By focusing policy updates on these prompts, AdaSelect effectively improves training stability. Extensive experiments across different T23D models and RL algorithms demonstrate that our framework consistently improves generation quality across multiple dimensions.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑