arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37670cs.AIcs.CEcs.CV

MeanFlowAdvantage:少步平均速度生成器的稳定奖励微调

MeanFlowAdvantage: Stable Reward Fine-Tuning for Few-Step Average-Velocity Generators

Haocheng Tang, Tianchi Xie, Xingqiao Lin

首次发表
浏览论文内容

中文总结 AI 辅助

针对平均速度生成器奖励微调不匹配问题,提出带符号优势加权最小二乘目标MeanFlowAdvantage,在少步采样下提升图像生成与DNA设计性能。

中文摘要 AI 辅助

MeanFlow通过预测区间平均速度实现高效的少步生成,但这种表示与奖励微调之间存在不匹配:现有的基于优势的目标通常定义在瞬时速度或等价的$x_0$空间预测上,而推理时直接使用学习到的平均速度映射。我们提出了MeanFlowAdvantage,一种用于平均速度生成器的带符号优势加权最小二乘目标。我们的关键构造使用共享的、分离的MeanFlow导数校正,在预测空间中表达奖励目标,同时使展开和参考正则化成为推理时部署的平均速度网络上的精确惩罚。由此产生的公式保留了MeanFlow原生的少步采样器,并为将奖励改进直接转移到部署的流映射提供了一种机制。在SD3.5-Medium上,MeanFlowAdvantage在匹配的四步MeanFlowNFT基线上改善了所有八项报告指标,并且仅用四次NFE,在八项指标中的六项上达到或超过了40步DiffusionNFT基线。相同的目标也适用于DNA启动子设计,它支持定义在流形上的生成器的无教师在线强化学习以及教师引导的奖励分级蒸馏,后者在比较配置中产生了最低的单步Sei轮廓MSE。

英文摘要

MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tuning: existing advantage-based objectives are typically defined on instantaneous velocities or equivalent $x_0$-space predictions, whereas inference directly uses the learned average-velocity map. We introduce MeanFlowAdvantage, a signed advantage-weighted least-squares objective for average-velocity generators. Our key construction uses a shared, detached MeanFlow derivative correction to express the reward objective in prediction space while making rollout and reference regularization exact penalties on the average-velocity network deployed at inference. The resulting formulation preserves MeanFlow's native few-step sampler and provides a direct mechanism for transferring reward improvements to the deployed flow map. On SD3.5-Medium, MeanFlowAdvantage improves all eight reported metrics over the matched four-step MeanFlowNFT baseline and, with only four NFEs, matches or exceeds the 40-step DiffusionNFT baseline on six of eight metrics. The same objective also transfers to DNA promoter design, where it supports both teacher-free on-policy RL for a generator defined on a manifold and teacher-guided reward-graded distillation, with the latter yielding the lowest one-step Sei profile MSE among the compared configurations.

发表机构

  • Northeastern University(东北大学)
  • Tsinghua University(清华大学)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

↑