发表机构
South China University of Technology; Nanyang Technological University(华南理工大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对有限演示预算下机器人操作泛化困难,提出ComManip,通过将策略专门化到舒适操作区域并重新定位基座,提升成功率约20个百分点,实现数据高效学习。
AI 中文摘要
训练机器人操作策略依赖于昂贵的机器人演示,使得大规模数据收集不切实际。同时,为了提高策略的泛化能力,现有方法通过在数据收集过程中改变物体放置、视角和机器人配置来寻求视觉观察的更大多样性。然而,在有限的演示预算下,这种策略迫使策略对多样的视觉观察进行建模,导致在相似的局部条件下学习可靠的观察-动作对应关系的监督不足。我们的研究表明,在这种策略下训练的策略,其任务成功率低于在紧凑、视觉和运动学稳定的区域内训练的策略。我们将这些稳定区域称为舒适操作区域。为了利用这一发现,我们提出了ComManip,一种将操作策略专门化到舒适操作区域的学习范式。在推理过程中,ComManip重新定位移动基座,直到检测到的目标中心进入从舒适区域演示估计的熟悉图像空间范围。然后执行相同的操作策略,从而实现对不同目标位置的有效操作。我们在多个操作任务、演示预算和策略家族(包括ACT、π₀.₅、RDT、OpenVLA-OFT和SmolVLA)上进行了大量实验。结果表明,在有限的演示预算下,ComManip在不同策略架构的大工作空间中,将任务成功率提高了大约20个百分点或更多,这表明将操作策略专门化到舒适区域提供了一种更数据高效的操作学习范式。
英文摘要
Training robot manipulation policies relies on costly robot demonstrations, making large-scale data collection impractical. Meanwhile, to improve policy generalization, existing approaches seek greater diversity in visual observations by varying object placements, viewpoints, and robot configurations during data collection. However, under a limited demonstration budget, this strategy forces the policy to model diverse visual observations, providing insufficient supervision to learn reliable observation-action correspondences under similar local conditions. Our study reveals that policies trained under this strategy achieve lower task success rates than those trained within a compact, visually and kinematically stable region. We refer to these stable regions as comfortable manipulation regions. To exploit this finding, we propose ComManip, a learning paradigm that specializes manipulation policies to comfortable manipulation regions. During inference, ComManip repositions the mobile base until the detected target center enters the familiar image-space range estimated from comfortable-region demonstrations. It then executes the same manipulation policy, enabling effective manipulation across diverse target locations. We conduct extensive experiments across multiple manipulation tasks, demonstration budgets, and policy families including ACT, $π_{0.5}$, RDT, OpenVLA-OFT, and SmolVLA. The results demonstrate that ComManip improves task success by roughly 20 percentage points or more across different policy architectures in large workspaces under limited demonstration budgets, suggesting that specializing manipulation policies to comfortable regions provides a more data-efficient learning paradigm for manipulation.