发表机构
South China University of Technology; Yuanwu Technology(华南理工大学; 元武科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RecastVLA通过流匹配策略中的自适应状态实现顺序操作,利用动作侧测试时训练提升成功率,在多个基准和真实任务中显著优于基线。
AI 中文摘要
顺序操作要求机器人跟踪已经发生的事情,即使当前场景不再显示这些信息。具有显式历史表示的策略使过去的交互可作为当前决策的上下文。我们探讨动作生成本身如何形成用于后续控制的持久状态。基于动作侧测试时训练,RecastVLA在流匹配视觉-语言-动作策略中维护一个自适应策略状态。该状态由共享的快速权重表示,并在整个动作生成过程中保持固定。深度特定的接口读取相同的状态,而跨深度和流评估的特征共同定义下一次策略调用的一个更新。后续动作损失通过区分先前的状态转换来训练初始化、接口和更新规则。在部署时,更新使用策略自身的动作生成特征,无需专家动作标签。在LIBERO、RoboTwin、RoboDojo和十二个真实机器人任务中,RecastVLA相对于未进行测试时训练的匹配策略提高了平均成功率,包括在RoboTwin Clean-to-Clean上提高了10.68个百分点。在受控的RoboTwin比较中,保留状态提高了成功率,共享设计比独立训练的逐层局部TTT高出2.58个百分点。
英文摘要
Sequential manipulation requires a robot to track what has already happened, even when the current scene no longer reveals it. Policies with explicit history representations make past interactions available as context for current decisions. We ask how action generation itself can form a persistent state for subsequent control. Building on action-side test-time training, RecastVLA maintains an adaptive policy state within a flow-matching vision-language-action policy. The state is represented by shared fast weights and remains fixed throughout action generation. Depth-specific interfaces read the same state, while features across depths and flow evaluations jointly define one update for the next policy call. Subsequent action losses train the initialization, interfaces, and update rule by differentiating through earlier state transitions. At deployment, updates use the policy's own action-generation features without expert action labels. Across LIBERO, RoboTwin, RoboDojo, and twelve real-robot tasks, RecastVLA improves mean success over a matched policy trained without test-time training, including 10.68 percentage points on RoboTwin Clean-to-Clean. In controlled RoboTwin comparisons, retaining state improves success, and the shared design exceeds independently trained layer-local TTT by 2.58 points.
Comments23 pages, 17 figures, 8 tables, including appendices