arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TraceFlow: 利用成功与失败轨迹引导冻结的流匹配机器人策略

TraceFlow: Guiding Frozen Flow-Matching Robot Policies with Success and Failure Traces

Jiaxuan Zhang, Ruizhe Liu, Yu Zhang, Yanchao Yang

arXiv 2609.20646首次发表:更新:

发表机构

The University of Hong Kong; Southern University of Science and Technology(香港大学; 南方科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TraceFlow利用成功与失败轨迹的引导场,在无需额外标签下修正冻结流匹配机器人策略,提升真实与仿真任务成功率并减少错误序列。

AI 中文摘要

视觉-语言-动作(VLA)策略配备流匹配动作专家,通过积分学习到的速度场来生成每个动作块(一个短命令序列);一旦其权重固定,先前回合的成功或失败无法改变当前生成的动作块。并行的测试时方法通过检索到的成功轨迹、学习到的评论家、验证器或动力学模型为冻结策略提供输入,但没有任何方法使用机器人自身的失败回合作为负面证据,且仅凭一个终端结果位。我们提出TraceFlow,一个进度对齐的引导场,将检索到的成功和失败回合的动作密度转化为对冻结流匹配动作专家的有界修正,每个回合仅使用一个终端结果位,无需其他标签。其TraceBank存储轨迹,即带终端标签的按时间排序的状态-动作记录,从目标任务训练轨迹开始,随后接纳部署机器人自身的回合。在有序的真实机器人打包任务中,基础策略按顺序完成50次试验中的21次,TraceFlow完成39次,且在一轮堆叠中无需任何权重更新即完成47次,错误序列回合从20降至0。在仿真中,增益具有选择性:在按套件选择的设置下,TraceFlow将RoboMemArena序列任务成功率从78.92%提升至91.50%,转移任务在堆叠第2轮从54.41%提升至62.00%,26任务聚合保持不变,计数和遮挡分别下降1.12和1.42个百分点,LIBERO-Plus(长)变化+1.27个百分点(p = 0.0733)。堆叠增益是有限的,每个分支在第10轮前达到峰值,且银行的成功-失败比率无法预测检索分配。

英文摘要

A vision-language-action (VLA) policy with a flow-matching action expert generates each action chunk (a short command sequence) by integrating a learned velocity field; once its weights are fixed, the success or failure of an earlier rollout cannot change the chunk generated now. Concurrent test-time methods give a frozen policy such an input from retrieved successes, a learned critic, a verifier, or a dynamics model, but none uses the robot's own failed rollouts as negative evidence with nothing but a terminal outcome bit. We introduce TraceFlow, a progress-aligned guidance field that turns the action densities of retrieved successful and failed rollouts into a bounded correction to a frozen flow-matching action expert, using one terminal outcome bit per rollout and no other label. Its TraceBank stores traces, time-ordered state-action records with a terminal label, starts from the target-task training traces, and later admits the deployed robot's own rollouts. On an ordered real-robot packing task the base completes 21 of 50 trials in order, TraceFlow 39, and one stacking round without any weight update 47, with wrong-sequence episodes falling from 20 to 0. In simulation the gain is selective: with per-suite selected settings, TraceFlow raises RoboMemArena Sequence from 78.92\% to 91.50\% task success and Transferring from 54.41\% to 62.00\% at stacking round 2, leaves the 26-task aggregate unchanged, lowers Counting and Occlusion by 1.12 and 1.42 points, and changes LIBERO-Plus (Long) by +1.27 points (p = 0.0733). Stacking gains are finite, every branch peaking before round ten, and the bank's success-to-failure ratio predicts no retrieval allocation. Project page: https://zhangjiaxuan-xuan.github.io/TraceFlow/

CommentsProject page: https://zhangjiaxuan-xuan.github.io/TraceFlow/ including code models and realworld-data

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑