QAMM:伴随均值流匹配用于少步离线强化学习
QAMM: Adjoint MeanFlow Matching for Few-Step Offline Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
QAMM将评论家伴随信号转化为对MeanFlow平均速度的监督,实现少步离线强化学习,在HumanoidMaze任务上两步生成动作并达到竞争性能。
中文摘要 AI 辅助
流策略能够建模丰富的动作分布,但其迭代采样限制了决策速度。伴随匹配利用评论家的动作梯度来改进流策略,而无需通过其采样轨迹进行反向传播,然而其监督是针对瞬时速度定义的。我们提出了QAMM,一种将评论家派生的伴随信号转化为对MeanFlow平均速度监督的方法。由此产生的策略直接学习有限区间传输,并通过少量网络评估生成动作。我们推导了伴随MeanFlow目标,指定了其梯度边界,并使用离线演员-评论家进行训练。在十个HumanoidMaze任务上,QAMM生成了有效的两步调用策略,并与强流策略基线相比取得了具有竞争力的性能。这些结果表明,基于伴随的Q优化可以与平均速度学习相结合,以获得具有少步动作生成能力的表达性离线策略。
英文摘要
Flow policies can model rich action distributions, but their iterative sampling limits decision speed. Adjoint matching uses the critic's action gradient to improve a flow policy without backpropagating through its sampling trajectory, yet its supervision is defined for instantaneous velocities. We propose QAMM, a method that turns the critic-derived adjoint signal into supervision for MeanFlow's average velocity. The resulting policy learns finite-interval transport directly and generates actions with few network evaluations. We derive the adjoint MeanFlow target, specify its gradient boundaries, and train it with an offline actor-critic. On ten HumanoidMaze tasks, QAMM produces effective two-call policies and achieves competitive performance against strong flow-policy baselines. These results show that adjoint-based Q optimization can be combined with average-velocity learning to obtain expressive offline policies with few-step action generation.
发表机构
- Fudan University(复旦大学)
- ShanghaiTech University(上海科技大学)
- Shanghai Innovation Institute(上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。