arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在线演化策略用于基于流匹配的视觉-语言-动作策略,通过自监督轨迹分布优化

Online Evolution Strategy for Flow-Matching VLA Policies via Self-Supervised Trajectory Distribution Optimization

Gongxin Yao, Yongsheng Zhao, Jiayin Deng, Deng Liang, Han Gao, Lei Zhao, Baoping Cheng

arXiv 2609.38855首次发表:更新:

发表机构

China Mobile (Hangzhou) Information Technology Co., Ltd.(中国移动(杭州)信息技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对流匹配VLA策略轨迹分布不良问题,提出在线演化策略Online-ES,在动作轨迹空间进行演化探索并自监督优化,无需价值模型即可达到强化微调效果。

AI 中文摘要

基于生成框架(如流匹配)的视觉-语言-动作(VLA)模型近期在机器人操作中取得了令人瞩目的性能。与确定性策略不同,流匹配使VLA模型能够学习条件动作轨迹分布,其中潜在噪声向量在同一任务场景下诱发不同的动作。然而,我们观察到这些分布往往形态不良,成功与失败的行为并存,而相当大的概率质量仍处于不利区域。为此,我们提出Online-ES,一种基于演化策略(ES)的流匹配VLA在线适应框架,通过交互反馈细化学习到的动作轨迹分布。我们的方法不在潜在噪声空间中进行剪枝,而是直接在动作轨迹空间中进行演化探索,其中由流匹配生成的多样化轨迹为适应提供候选解。通过对采样轨迹进行扰动并评估其执行结果,我们推导出一个自监督的均方误差目标,将演化方向从轨迹空间转移到模型参数空间。数学上,我们证明了所提出的目标是最优演化方向的无偏估计量。此外,我们还纳入失败经验作为负反馈以正则化演化方向,使策略远离先前探索过的失败区域。在仿真和真实环境中的实验表明,Online-ES实现了与强化微调相当的策略改进,而无需学习价值模型或计算优势。

英文摘要

Vision-Language-Action (VLA) models based on generative frameworks, such as Flow Matching, have recently achieved impressive performance in robotic manipulation. Unlike deterministic policies, Flow Matching enables VLA models to learn conditional action trajectory distributions, where latent noise vectors induce different actions under the same task scenario. However, we observe that these distributions are often ill-formed, with successful and failed behaviors coexisting while considerable probability mass remains in unfavorable regions. To this end, we propose Online-ES, an online adaptation framework for Flow Matching VLAs based on Evolution Strategy (ES), which refines the learned action trajectory distribution through interaction feedback. Instead of pruning the latent noise space, our method performs evolutionary exploration directly in the action trajectory space, where diverse trajectories generated by Flow Matching provide candidate solutions for adaptation. By perturbing sampled trajectories and evaluating their execution outcomes, we derive a self-supervised MSE objective that transfers the evolution direction from trajectory space into model parameter space. Mathematically, we prove that the proposed objective provides an unbiased estimator of the optimal evolution direction. Moreover, we also incorporate failure experiences as negative feedback to regularize the evolution direction, steering the policy away from previously explored failure regions. Experiments in both simulation and real-world environments demonstrate that Online-ES achieves policy improvement comparable to reinforcement fine-tuning, without learning a value model or computing advantages.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑