更少步骤,更好动作:重新思考VLA策略的流匹配推理
Fewer Steps, Better Actions: Rethinking Flow-Matching Inference for VLA Policies
浏览论文内容
中文总结 AI 辅助
提出Coda,通过轻量级Transformer学习端点校正,替代部分积分步骤,在降低延迟的同时提升VLA策略的成功率,验证了校正作为额外积分有效替代方案的价值。
中文摘要 AI 辅助
基于流匹配的视觉-语言-动作(VLA)策略通过反复评估动作专家来生成动作块。增加积分步数会提高推理成本,但不一定能改善闭环成功率。我们提出了Coda,它将部分积分预算重新分配给一个学习到的端点校正。一个冻结的策略首先完成几步从噪声到动作的轨迹;然后一个轻量级Transformer利用候选动作、源噪声和共享的观察前缀缓存来预测一个由示范监督的残差。只有校正器被训练。在50个RoboTwin Easy任务上,五步Coda将成功率从71.64%提高到74.68%,相比匹配的五步基线,同时相对于默认的十步策略将前向延迟降低了30.2%。一个两步配置实现了71.88%的成功率,并获得了2.12倍的加速。一个独立的13任务对照显示,在几乎相同的延迟下获得了5.69个百分点的提升,支持将校正作为额外积分的有效替代方案。同样的设计也改进了冻结的官方SmolVLA,将两步成功率从60.8%提高到69.4%。这些结果表明,端点校正改善了冻结流匹配策略的质量-延迟权衡。
英文摘要
Vision-language-action (VLA) policies based on flow matching generate action chunks through repeated evaluations of an action expert. Increasing the number of integration steps raises inference cost, but does not necessarily improve closed-loop success. We propose Coda, which reallocates part of this integration budget to a single learned endpoint correction. A frozen policy first completes a few-step noise-to-action trajectory; a lightweight Transformer then predicts a demonstration-supervised residual using the candidate action, source noise, and shared observation-prefix cache. Only the corrector is trained. On 50 RoboTwin Easy tasks, five-step Coda improves success from 71.64% to 74.68% over the matched five-step baseline, while reducing forward latency by 30.2% relative to the default ten-step policy. A two-step configuration achieves 71.88% success with a 2.12$\times$ speedup. An independent 13-task control shows a 5.69-percentage-point gain at nearly equal latency, supporting correction as an effective alternative to additional integration. The same design also improves frozen official SmolVLA, raising two-step success from 60.8% to 69.4%. These results show that endpoint correction improves the quality-latency trade-off of frozen flow-matching policies.
发表机构
- Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
- Zhongke Haichuan Intelligent(中科海川智能)
机构由 AI 辅助整理,请以论文原文为准。