发表机构
Jilin University; Peking University; Chinese Academy of Sciences; University of Birmingham(吉林大学; 北京大学; 中国科学院; 伯明翰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对扁平物体机器人操作泛化能力有限的问题,提出含策略生成器与执行模块的统一框架,并构建FlatLab仿真基准,方法泛化效果优于现有基线。
AI 中文摘要
扁平物体的机器人操作存在诸多挑战,原因在于其不可抓取构型以及物体几何形状和材料的巨大差异。现有方法依赖启发式预操作,且常在封闭场景中评估,泛化能力有限。我们提出一种统一框架,将操作解耦为策略生成器和动作执行模块:策略生成器通过模拟数据转换与对比学习,学习以策略为中心、与物体无关的表征,进而从物体点云预测合适的操作策略;执行模块则基于预测的策略,将长程操作分解为可复用的动作原语,并动态组合这些原语以生成稳定轨迹。为实现系统性评估,我们引入FlatLab——一个用于扁平物体机器人操作的综合仿真基准测试,其提供多样刚性与可变形扁平物体的高保真物理仿真、自动化多模态数据采集,以及标准化任务定义与评估协议。在FlatLab中开展的实验表明,我们的方法能有效泛化至未见过的物体及类别,性能优于现有基线方法。项目页面与代码可在此httpsURL获取。
英文摘要
Robotic manipulation of flat objects is challenging due to the ungraspable configurations and strong variations in object geometry and material. Existing methods rely on heuristic pre-manipulation and are often evaluated in closed settings with limited generalization. We propose a unified framework that decouples the manipulation into a strategy generator and an action execution module. The strategy generator predicts appropriate manipulation strategies from object point clouds by learning strategy-centric, object-invariant representations via simulated data transformation and contrastive learning. Conditioned on the predicted strategy, the execution module decomposes long-horizon manipulation into reusable action primitives and dynamically composes them to generate stable trajectories. To enable systematic evaluation, we introduce FlatLab, a comprehensive simulation benchmark for robotic flat object manipulation. FlatLab provides high-fidelity physical simulation of diverse rigid and deformable flat objects, automated multi-modal data collection, and standardized task definitions and evaluation protocols. Experiments conducted in FlatLab demonstrate that our approach generalizes effectively to unseen objects and categories, outperforming existing baselines. The project page and the code are provided at https://flatlab-web.github.io/.
CommentsThis paper is accepted to ICML 2026