AI 中文总结
ChainVLA 是 12 亿参数的 VLA 策略,通过联合可修订的执行状态串联查询,结合进展上下文与运动尾部,在长 horizon 操纵任务中大幅提升成功率, ablation 验证了两个组件的关键作用。
AI 中文摘要
人类执行长 horizon 操纵任务时,会保留早期动作建立的知识,同时持续调整正在进行的运动。相比之下,分块动作的视觉-语言-动作(VLA)策略会在每次查询时根据当前输入重复重新规划。现有方法要么通过记忆保留长期任务证据,要么通过动作复用和集成保留短期运动,但跨查询的交接仍不完整。我们提出 ChainVLA,这是一个拥有 12 亿参数的 VLA 策略,它通过联合且可修订的执行状态串联连续查询。进展上下文(Progress Context)结合循环工作状态(Working State)与稀疏事件记忆,承载由观测得到的任务进展;运动尾部(Motion Tail)则将前一次预测未执行的延续部分输入状态构建和动作生成过程。这两个组件共同为解码器提供条件,使其在最新观测下重新生成每个动作 horizon,同时承载的状态可指导下一次预测而不将其固定。ChainVLA 在 RMBench 上达到 62.8% 的平均成功率,在四个 LIBERO 套件上达到 98.8% 的成功率;移除 Motion Tail 会使 RMBench 成功率降至 11.2%,移除 Progress Context 则降至 3.0%。这些不对称的 ablation 结果表明,运动连续性有助于保留观测流,而任务进展正是从该观测流中推断得出的。
英文摘要
Humans perform long-horizon manipulation by retaining knowledge of what earlier actions have established while continuously adapting the motion underway. By contrast, action-chunked vision-language-action (VLA) policies repeatedly replan from the current input at each query. Existing methods preserve either long-term task evidence through memory or short-term motion through action reuse and ensembling, leaving the cross-query handoff incomplete. We introduce ChainVLA, a 1.2B-parameter VLA policy that chains successive queries through a joint and revisable execution state. Progress Context combines a recurrent Working State with sparse event memory to carry observation-derived task progress, while Motion Tail feeds the preceding prediction's unexecuted continuation into state construction and action generation. Together, the two components condition a decoder that regenerates each action horizon under the latest observation, allowing the carried state to guide the next prediction without fixing it. ChainVLA reaches 62.8% average success on RMBench and 98.8% across four LIBERO suites, while removing Motion Tail or Progress Context reduces RMBench success to 11.2% and 3.0%, respectively. These asymmetric ablations are consistent with motion continuity helping preserve the observation stream from which task progress is inferred.
Comments13 pages (9 main + 4 appendix), 4 figures. Project page: https://muqy1818.github.io/chainvla-web/