发表机构
Microsoft Research Asia - Tokyo; Graduate School of Information Science and Technology, the University of Osaka(微软亚洲研究院(东京); 大阪大学信息科学与技术研究生院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散/流匹配动作模型推理延迟高的问题,提出无教师ESP策略,用能量分数训练动作头,一步生成动作块,在仿真和真实任务中成功率相当且延迟显著降低。
AI 中文摘要
基于扩散和流匹配的生成式动作模型在视觉-语言-动作(VLA)策略中日益被采用,因为它们能够捕捉多样化的行为,包括在同一观察和指令下的多个有效动作序列。然而,其迭代采样过程需要重复的网络评估来生成每个动作块,增加了闭环控制中的推理延迟。我们提出ESP(能量分数策略),一种无教师的方法,将策略上下文和噪声直接映射到单个网络评估中的动作块。ESP使用能量分数而非均方误差来训练动作头。平方误差回归针对条件均值,而能量分数是严格适当的:其期望值由目标分布唯一最小化。这为学习多模态动作分布提供了原则性的目标,无需迭代采样,当模型能够表示目标分布时,在总体最优处可精确恢复。在仿真和真实世界操作任务上的实验表明,与流匹配基线相比,任务成功率相当,但动作生成延迟显著降低。这些结果支持直接分布学习作为迭代生成式机器人策略的高效替代方案。
英文摘要
Generative action models based on diffusion and flow matching have been increasingly adopted in vision-language-action (VLA) policies for their ability to capture diverse behaviors, including multiple valid action sequences under the same observation and instruction. Their iterative sampling procedures, however, require repeated network evaluations to generate each action chunk, increasing inference latency in closed-loop control. We propose ESP (Energy-Score Policy), a teacher-free approach that maps policy context and noise directly to an action chunk in a single network evaluation. ESP trains the action head with the energy score rather than mean squared error. Whereas squared-error regression targets the conditional mean, the energy score is strictly proper: its expected value is uniquely minimized by the target distribution. This provides a principled objective for learning multimodal action distributions without iterative sampling, with exact recovery at the population optimum when the model can represent the target distribution. Experiments on both simulation and real-world manipulation tasks demonstrate competitive task success with substantially lower action-generation latency than the flow matching baseline. These results support direct distributional learning as an efficient alternative to iterative generative robot policies.