时间策略:用于机器人演示学习的历史初始化动作生成
Temporal Policy: History-Initialized Action Generation for Robotic Learning from Demonstration
浏览论文内容
中文总结 AI 辅助
本文提出Temporal Policy生成框架,将动作生成为时间耦合传输问题,在降低近一个数量级传输成本的同时匹配基线成功率,实现高频闭环控制。
中文摘要 AI 辅助
标准扩散模型和流匹配模型依赖于无信息高斯先验的独立耦合,被迫学习复杂且高成本的向量场以到达物理动作空间。生成模型擅长捕捉机器人演示学习(LfD)中的多模态行为,但通常存在推理成本高的问题。本文提出Temporal Policy,一种基于随机插值的生成框架,将动作生成为时间耦合传输问题。通过在机器人近期历史上初始化生成流,我们明确地将过去状态与未来动作序列耦合。这种依赖数据的耦合降低了传输成本,产生了平滑的向量场。我们在视觉运动模拟基准和物理Barrett WAM 2×7自由度遥操作平台上验证了Temporal Policy。与噪声初始化基线相比,我们的方法将传输成本降低了近一个数量级,在单个NVIDIA RTX 4080上实现了19.1毫秒的推理延迟。重要的是,在匹配最先进基线成功率的同时,实现了这些几何和计算效率。这种简化的传输几何绕过了独立高斯先验的计算瓶颈,有助于实现高频闭环控制。代码可在此https URL获取。
英文摘要
By relying on independent couplings from uninformative Gaussian priors, standard diffusion and flow matching models are forced to learn complex, high-cost vector fields to reach the physical action space. Generative models excel at capturing multimodal behaviors for robotic Learning from Demonstration (LfD), but often suffer from high inference cost. This paper introduces Temporal Policy, a generative framework based on stochastic interpolants that formulates action generation as a temporally coupled transport problem. By initializing the generative flow at the robot's recent history, we explicitly couple past states to future action sequences. This data-dependent coupling reduces transport cost and produces straight vector fields. We validate Temporal Policy across visuomotor simulation benchmarks and on a physical Barrett WAM 2x 7DoF teleoperation platform. Our approach reduces transport costs by nearly an order of magnitude compared to noise-initialized baselines, achieving a 19.1 ms inference latency on a single NVIDIA RTX 4080. Crucially, these geometric and computational efficiencies are achieved while matching the success rates of state-of-the-art baselines. This simplified transport geometry bypasses the computational bottleneck of independent Gaussian priors, helping enable high-frequency, closed-loop control. The code is publicly available at https://github.com/dmiller12/TemporalPolicy.
发表机构
- University of Alberta(阿尔伯塔大学)
机构由 AI 辅助整理,请以论文原文为准。