发表机构
Purdue University; University of Oxford(普渡大学; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLA异步执行导致动作分布偏移、反应性下降的问题,提出递归流场蒸馏与提议-解决机制,对齐异步与原始动作分布,在LIBERO持平成功率,在RoboMimic保留约80%成功率。
AI 中文摘要
通用机器人策略(如视觉-语言-动作模型,VLA)已实现显著的泛化能力,但其推理延迟可能与实时控制的需求相冲突。异步执行通过在机器人执行前一个动作序列的同时预测下一个动作序列,避免了动作块之间的停顿。本文研究异步执行是否产生与原始VLA相同的动作分布。我们发现,对于非马尔可夫演示,异步执行可能产生根本不同的动作分布,从而限制策略的反应性。在我们的方法中,我们通过两种互补机制,将异步产生的动作分布与原始VLA的动作分布对齐,以恢复这种反应性。首先,递归流场蒸馏利用VLA的动作生成流来训练异步策略。我们从理论上刻画了学习到的分布,并通过实验表明,我们的异步策略能够生成原始VLA几乎全部的动作范围,而现有的异步方法仅能恢复该范围的一小部分。其次,提议-解决机制异步准备多个动作序列,并使用最新观测基于它们在VLA动作分布下的似然性的轻量近似进行选择。我们的最终方法在LIBERO上匹配原始VLA的成功率,并在RoboMimic上保留约80%的成功率,比现有异步方法高出约30个百分点。
英文摘要
Generalist robot policies such as vision-language-action models (VLAs) have achieved remarkable generalization, but their inference delays can conflict with the demands of real-time control. Asynchronous execution avoids pauses between action chunks by predicting the next sequence of actions while the robot carries out the previous one. In this paper, we study whether asynchronous execution produces the same action distribution as the original VLA. We find that, for non-Markovian demonstrations, asynchronous execution can produce a fundamentally different action distribution, which can limit the policy's reactivity. In our method, we seek to restore this reactivity by aligning the asynchronously produced action distribution with that of the original VLA through two complementary mechanisms. First, Recursive Flow-Field Distillation trains the asynchronous policy using the VLA's action-generation flow. We characterize the learned distribution theoretically and show experimentally that our asynchronous policy can generate nearly the full range of actions the original VLA would produce, while existing asynchronous methods recover only a fraction of that range. Second, Propose-Resolve prepares multiple action sequences asynchronously and uses the latest observation to select among them based on a lightweight approximation of their likelihood under the VLA's action distribution. Our resulting method matches the original VLA's success on LIBERO and retains about 80% of its success on RoboMimic, about 30 percentage points more than existing asynchronous methods.