arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13425cs.LGcs.AI

ReCAST:面向在线扩散强化学习的跨时间步奖励信用分配

ReCAST: Reward Credit Assignment across Timesteps for Online Diffusion Reinforcement

  • University of California, Los Angeles(加州大学洛杉矶分校)
  • Arena AI

机构由 AI 辅助整理,请以论文原文为准。

Yihang Chen, Yuanhao Ban, Kuei-Chun Kao, Cho-Jui Hsieh

AI总结:

ReCAST提出按时间步和奖励的权重矩阵,根据Rényi可区分性增益分配信用,以区分用户偏好与奖励信息量,在扩散奖励微调中提升训练及泛化性能。

AI中文摘要:

使用多个奖励训练扩散模型需要区分用户偏好与奖励信息量。用户偏好决定了每个奖励对整体目标的贡献程度;奖励信息量决定了在去噪过程中其反馈何时有用。有些奖励在全局结构一出现时就能有意义地评估样本,但其他奖励只有在样本接近干净时才变得有信息量。为了同时解决这两个问题,我们提出了ReCAST(跨时间步奖励信用分配),据我们所知,这是扩散奖励微调中首个针对每个奖励、依赖时间步的信用分配方法。ReCAST通过一个按奖励和时间步的权重矩阵$W$将用户偏好与时间分配分离,该矩阵的行和与用户指定的奖励预算$\lambda$匹配,而列和相等,为每个去噪步骤分配相同的总权重。在这些边际约束下,ReCAST根据每个奖励的信息量分配权重,信息量通过其在每个步骤的Rényi可区分性增益量化。这些增益可伸缩到奖励诱导的正策略与当前策略之间的总可区分性,为时间信用分配提供了基础。我们通过在两个不同的四奖励设置下训练SD3.5-Medium来评估ReCAST,每个设置跨越五个奖励预算$\lambda$。ReCAST在一个设置中提高了训练奖励,在另一个设置中与之匹配,在两个设置中都提高了每个保留的评判者,并且被独立的LLM-as-a-Judge所偏好。这些结果共同表明,ReCAST产生的改进超越了训练奖励,并支持其核心原则:在反馈最具信息量的去噪时间步为每个奖励分配更大的权重。

英文摘要:

Training diffusion models with multiple rewards requires distinguishing user preference from reward informativeness. User preference determines how much each reward should contribute to the overall objective; reward informativeness determines when its feedback is useful during denoising. Some rewards can meaningfully evaluate a sample as soon as global structure emerges, but others become informative only when the sample is nearly clean. To address both questions jointly, we propose ReCAST (Reward Credit ASsignment across T}imesteps), the first method, to our knowledge, for per-reward, timestep-dependent credit assignment in diffusion reward fine-tuning. ReCAST separates user preferences from temporal allocation through a reward-by-timestep weight matrix $W$, whose row sums match the user-specified reward budgets $λ$, while its column sums are equal, assigning the same total weight to each denoising step. Under these marginal constraints, ReCAST allocates weight according to each reward's informativeness, quantified by its Rényi discriminability gain at each step. These gains telescope to the total discriminability between the reward-induced positive policy and the current policy, providing a basis for temporal credit assignment. We evaluate ReCAST by training SD3.5-Medium under two distinct four-reward settings, each across five reward budgets $λ$. ReCAST improves the training rewards in one setting and matches them in the other, improves every held-out judge in both, and is preferred by an independent LLM-as-a-Judge. Together, these results show that ReCAST yields improvements that generalize beyond the training rewards and support its core principle: assigning each reward greater weight at the denoising timesteps where its feedback is most informative.

↑