arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34497cs.LG

QAMM:伴随均值流匹配用于少步离线强化学习

QAMM: Adjoint MeanFlow Matching for Few-Step Offline Reinforcement Learning

Yuehu Gong, Shutong Ding, Mokai Pan, Yimiao Zhou, Jiashu Hou, Ye Shi, Yanwei Fu

首次发表
浏览论文内容

中文总结 AI 辅助

QAMM将评论家伴随信号转化为对MeanFlow平均速度的监督,实现少步离线强化学习,在HumanoidMaze任务上两步生成动作并达到竞争性能。

中文摘要 AI 辅助

流策略能够建模丰富的动作分布,但其迭代采样限制了决策速度。伴随匹配利用评论家的动作梯度来改进流策略,而无需通过其采样轨迹进行反向传播,然而其监督是针对瞬时速度定义的。我们提出了QAMM,一种将评论家派生的伴随信号转化为对MeanFlow平均速度监督的方法。由此产生的策略直接学习有限区间传输,并通过少量网络评估生成动作。我们推导了伴随MeanFlow目标,指定了其梯度边界,并使用离线演员-评论家进行训练。在十个HumanoidMaze任务上,QAMM生成了有效的两步调用策略,并与强流策略基线相比取得了具有竞争力的性能。这些结果表明,基于伴随的Q优化可以与平均速度学习相结合,以获得具有少步动作生成能力的表达性离线策略。

英文摘要

Flow policies can model rich action distributions, but their iterative sampling limits decision speed. Adjoint matching uses the critic's action gradient to improve a flow policy without backpropagating through its sampling trajectory, yet its supervision is defined for instantaneous velocities. We propose QAMM, a method that turns the critic-derived adjoint signal into supervision for MeanFlow's average velocity. The resulting policy learns finite-interval transport directly and generates actions with few network evaluations. We derive the adjoint MeanFlow target, specify its gradient boundaries, and train it with an offline actor-critic. On ten HumanoidMaze tasks, QAMM produces effective two-call policies and achieves competitive performance against strong flow-policy baselines. These results show that adjoint-based Q optimization can be combined with average-velocity learning to obtain expressive offline policies with few-step action generation.

发表机构

  • Fudan University(复旦大学)
  • ShanghaiTech University(上海科技大学)
  • Shanghai Innovation Institute(上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑