arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22041cs.LG

$λ$-Controlled GRPO:将流匹配比率不稳定性转化为预算化资源

$λ$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource

Yufeng Wang, Parivesh Priye, Meeshawn Marathe, Ramit Pahwa

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出$λ$-Controlled GRPO,将流匹配强化学习中的路径方差作为可预算资源,按预测规律校准重要性比率并分配梯度,提升文本准确性和偏好奖励。

中文摘要 AI 辅助

强化学习越来越多地被用于使图像生成器与奖励信号对齐,而Flow-GRPO最近通过将去噪采样器视为可从奖励反馈中优化的随机策略,将该范式扩展到了流匹配模型。在这种设置下的训练在多步去噪方面具有特定的不稳定性:策略更新在去噪步骤间系统性变化,重要性比率漂移低于1,变得越来越分散,以不同速率被裁剪,并在训练后期留下更少的可用样本。先前的工作将这些效应视为独立的失败模式,并用手工调优的稳定器分别处理。我们反而表明,它们源于一个单步量,我们称之为路径方差。该量由采样器的高斯转移核精确确定,并可在训练期间廉价估计。这重新将不稳定性定义为一种可测量和预算化的资源,而非需要修复的症状集合。我们的方法$λ$-Controlled GRPO根据这一预测规律而非嘈杂的经验统计来校准重要性比率行为,并根据预测成本在去噪步骤间分配梯度努力。控制更新的两个尺度由标准策略选择固定,而非作为自由调优参数引入。在两种奖励设置下的文生图模型上,即通过光学字符识别评分的困难目标文本渲染和通过偏好模型评分的人类偏好匹配,$λ$-Controlled GRPO在文本准确性和偏好奖励方面均优于最强的经验稳定器。它还将后期步骤的路径方差保持在预期预算内,而基线恰恰在系统性超调之处。结果是Flow-GRPO更新由其自身转移规律校准,而非在不稳定性出现后进行稳定化。

英文摘要

Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback. Training in this setting is unstable in a way specific to multi-step denoising: the policy update changes systematically across denoising steps, with importance ratios drifting below one, becoming increasingly dispersed, clipping at different rates, and leaving fewer usable samples late in training. Prior work treats these effects as separate failure modes and addresses each with a hand-tuned stabilizer. We show instead that they arise from a single per-step quantity, which we call path variance. This quantity is determined exactly by the sampler's Gaussian transition kernel and can be estimated cheaply during training. This reframes instability as a resource that can be measured and budgeted rather than a collection of symptoms to repair. Our method, $λ$-Controlled GRPO, calibrates importance-ratio behavior from this predicted law rather than from noisy empirical statistics, and allocates gradient effort across denoising steps according to their predicted cost. The two scales governing the update are fixed by standard policy choices rather than introduced as free tuning parameters. On a text-to-image model under two reward settings, rendering difficult target text scored by optical character recognition and matching human preferences scored by a preference model, $λ$-Controlled GRPO improves both text accuracy and preference reward over the strongest empirical stabilizer. It also keeps late-step path variance within its intended budget, precisely where the baseline systematically overshoots. The result is a Flow-GRPO update calibrated by its own transition law rather than stabilized after instability appears.

发表机构

  • Stony Brook University(石溪大学)
  • Rivian and Volkswagen Group Technologies(Rivian和大众集团科技公司)

机构由 AI 辅助整理,请以论文原文为准。

↑