发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究通用操作策略反应性和实时性问题,提出$\pi\mathbf{R}^2$方法,将条件分快速和慢速通道,采用延迟自适应流调度,能在保留大型主干等的同时,让策略实时反应并提高成功率。
AI 中文摘要
通用操作策略越来越多地采用基于大型预训练主干构建的动作分块流策略形式。这些块开环运行,政策无法对执行中途到达的感官输入做出反应,牺牲了反应性。更频繁地重新规划可以恢复它,但感知到行动的管道(一个大主干加上多个去噪步骤)太慢:这种延迟禁止频繁重新规划,并使已提交的行动过时,使得这些策略不适用于动态闭环控制。我们提出了$\pi\mathbf{R}^2$,它使这些策略具有反应性和实时性,同时保留大型主干、富有表现力的多模态策略和多动作预测。基于扩散强迫的逐位置噪声调度,$\pi\mathbf{R}^2$贡献了两个想法。首先,它将条件分为一个快速通道(本体感觉,每滴答更新一次)和一个异步更新的慢速通道(视觉语言特征),因此策略在一个块内对本体感觉做出反应,同时容忍过时的视觉。其次,延迟自适应流调度将飞行中的动作视为修复条件,并在每次调用的一个去噪步骤中发出动作,使一个经过训练的模型能够适应不同的硬件延迟。对现有架构只需进行最小的修改,$\pi\mathbf{R}^2$就可以从预训练策略中进行微调:在真实的xArm6+XHand平台上应用于GR00T-N1.7时,它的闭环重新规划速度比基本策略快约4倍(在A5000 GPU上约为25Hz),每40毫秒根据新观察进行一次动作。在模拟和实际操作任务中,$\pi\mathbf{R}^2$比最强基线在模拟中成功率提高高达23%,在现实世界中提高30%。项目页面:此https URL
英文摘要
Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing \emph{reactivity}. Replanning more often would restore it, but the perception-to-action pipeline (a large backbone plus multiple denoising steps) is too slow: this \emph{latency} forbids frequent replanning and leaves committed actions stale, making such policies ill-suited for dynamic, closed-loop control. We present $π\mathbf{R}^2$, which makes these policies reactive and real-time while retaining large backbones, expressive multi-modal policies, and multi-action prediction. Built on the per-position noise schedule of diffusion forcing, $π\mathbf{R}^2$ contributes two ideas. First, it splits conditioning into a fast channel (proprioception, fresh every tick) and an asynchronously updated slow channel (vision-language features), so the policy reacts to proprioception within a chunk while tolerating stale vision. Second, a latency-adaptive flow schedule treats in-flight actions as inpainting conditioning and emits actions in one denoising step per call, letting one trained model adapt to varying hardware latency. Requiring minimal modification to existing architectures, $π\mathbf{R}^2$ can be finetuned from a pretrained policy: applied to GR00T-N1.7 on a real xArm6+XHand platform, it replans closed-loop roughly $4\times$ faster than the base policy (~$25$Hz on an A5000 GPU), acting on a fresh observation every $40$ms. Across simulation and real-world manipulation tasks, $π\mathbf{R}^2$ improves the success rate by up to $23\%$ in simulation and $30\%$ in the real world over the strongest baseline. Project page: https://pi-r2-flow.github.io/
CommentsPreprint(20 pages). Under Review