arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32665cs.LGcs.AIcs.CV

时间步加权:有效的基于ELBO的流匹配强化学习中的隐藏关键

Timestep Weighting: A Hidden Key to Effective ELBO-Based Flow-Matching RL

Qinwei Ma, Jingzhe Shi, Simin Fan, Ling Li, Alex Lamb

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨了基于ELBO的流匹配强化学习中时间步加权的影响,发现有效加权依赖于奖励景观和学习阶段,并提出了改进性能的静态、预算及动态加权策略。

中文摘要 AI 辅助

基于ELBO的强化学习提供了一种与采样器无关的方法,用于通过奖励反馈微调流匹配模型。在基于ELBO的强化学习中,时间步加权对性能有重大影响,并且它提供了一个统一的视角(如本工作所示)来理解先前工作中启发式选择的预测损失,然而这一领域研究不足,且常常被选择继承预训练配置。我们研究了基于ELBO的强化学习中时间步加权的影响和动态。我们表明,有效的加权既取决于奖励景观,也取决于学习阶段。(1)通过在受控的CIFAR图像生成实验以及机器人技术上的补充实验,我们研究了加权如何影响跨噪声水平的奖励驱动更新。(2)通过梯度分析,我们揭示了跨任务噪声协调的不同模式及其在训练过程中的演变。这些发现支持了一个假设:有用的加权取决于策略当前行为与奖励所偏好的行为之间的差距。(3)在此分析的指导下,我们研究了简单的静态加权、预算配置文件选择和动态调度,这些方法在传统目标选择之外提高了性能。我们的结果确立了时间步加权作为流匹配强化学习中的一个重要设计选择,并激励了进一步研究在学习过程中选择和调整它的方法。

英文摘要

ELBO-based reinforcement learning offers a sampler-agnostic approach to fine-tuning flow matching models with reward feedback. Timestep weighting in ELBO-based RL has large impact on performance, and it also provides a unified view (as we show in this work) to understand prediction losses heuristically chosen in prior work, yet it remains under-researched and is often chosen to inherit pretrain configs. We investigate impacts and dynamics of timestep weighting in ELBO-based RL. We show that effective weighting depends on both the reward landscape and stage of learning. (1) Through experiments on controlled CIFAR image generation, complemented by robotics, we investigate how weighting impacts reward-driven updates across noise levels. (2) Through gradient analysis, we reveal distinct patterns of cross-noise coordination across tasks and their evolution during training. These findings motivate the hypothesis that useful weighting depends on the gap between the policy's current behavior and the behavior favored by the reward. (3) Guided by this analysis, we study simple static weighting, budgeted profile selection, and dynamic schedules that improve performance beyond conventional target choices. Our results establish timestep weighting as an important design choice for flow-matching RL and motivate further research into methods that choose and adapt it throughout learning.

补充信息

↑