arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过黎曼均值流在动作流形上实现更快的视觉运动策略学习

Faster Visuomotor Policy Learning on Action Manifolds via Riemannian MeanFlow

S. Talha Bukhari, Austin Garrett, Yi Wei, Ruiqi Ni, Zachary Kingston, Aniket Bera

arXiv 2609.30127首次发表:更新:

发表机构

Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出黎曼均值流策略(RMFP),在动作流形上学习条件流映射,通过流映射一致性目标实现单次网络评估生成动作序列,在多个基准和真实任务上以更低采样成本达到竞争性能。

AI 中文摘要

视觉运动策略学习从原始感官观测到机器人动作序列的直接映射。基于扩散和流匹配的策略以端到端的方式捕获动作序列上的多模态分布。这种表达能力的代价是,在动作生成过程中需要对学习到的向量场进行多步数值积分,这可能既昂贵又耗时,阻碍了机器人应用所需的高速控制率。此外,机器人动作序列通常定义在光滑可微的流形上,要求学习到的策略尊重机器人动作空间的内在几何结构。在此,我们提出了黎曼均值流策略(RMFP),该策略学习机器人动作流形上概率路径的条件流映射。我们的公式采用了一个流映射一致性目标,该目标通过黎曼条件流匹配锚点基于数据进行约束。流映射一致性条件训练稳定,并将学习到的模型限制为有限时间传输,从而只需一次网络函数评估即可生成流形上的动作序列。我们在球形LASA和Push-T基准测试、Robomimic套件的Tool Hang和Transport任务以及具有流形约束动作生成的Franka Kitchen任务上展示了结果,并证明RMFP以更低的采样成本达到了与先前工作相当的性能。我们还将RMFP应用于真实世界的机器人操作任务,以展示在物理世界中传感器测量不完美的情况下快速动作生成的能力。

英文摘要

Visuomotor policies learn a direct map from raw sensory observations to robot action sequences. Policies based on Diffusion and Flow Matching capture the multimodal distribution over action sequences in an end-to-end manner. This expressivity comes at the cost of multi-step numerical integration of the learned vector field for action generation, which can be expensive and time-consuming, impeding fast control rates required in robotics applications. Furthermore, robot action sequences are usually defined on a smooth, differentiable manifold, requiring that the learned policy respects the intrinsic geometry of the robot's action space. Here, we present Riemannian MeanFlow Policy (RMFP), which learns the conditioned flow map of the probability path on the robot action manifold. Our formulation employs a flow map consistency objective grounded in the data by a Riemannian Conditional Flow Matching anchor. The flow map consistency condition is stable to train and constrains the learned model to finite-time transport, which yields on-manifold action sequence generation with as few as one network function evaluation. We present results on the spherical LASA and Push-T benchmarks, on the Tool Hang and Transport tasks of the Robomimic suite, and on the Franka Kitchen task with manifold-constrained action generation, and demonstrate that RMFP attains performance competitive with prior work at a lower sampling cost. We also employ RMFP on a real-world robotic manipulation task to demonstrate fast action generation under imperfect sensor measurements in the physical world.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑