发表机构
NLPR, MAIS, Institute of Automation, Chinese Academy of Sciences; School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院自动化研究所; 中国科学院大学; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AIM框架将真实到仿真软体仿真视为局部-全局交互建模问题,利用运动历史和几何调整粒子关系,并通过统一控制点接口和多步自回归训练,在PhysTwin和PGND上显著降低预测误差,支持跨场景零样本迁移。
AI 中文摘要
可变形物体操作对于机器人任务(如折叠衣物和处理食物)至关重要,在这些任务中,机器人必须控制形状变化以及物体运动。预测性软体仿真通过预测外部交互下的变形来支持这些任务。然而,空间邻域可能错误地表示变形依赖关系,引入局部误差,这些误差会在连续预测中累积。针对单个场景拟合的模型还必须适应物体几何形状和操作条件的变化。在这项工作中,我们提出了AIM,一个自适应交互建模框架,将真实到仿真软体仿真视为一个局部-全局交互建模问题。AIM利用运动历史和几何形状,在当前空间邻域和保留连接上调整粒子关系,而几何条件化的全局通信协调物体范围内的响应。统一的运动学控制点接口表示不同的操作配置,多步自回归监督训练模型基于其自身预测的轨迹。在PhysTwin和PGND上的实验展示了改进的运动准确性和视觉保真度,相对于PhysTwin,未来预测跟踪误差降低了20.0%,相对于PGND,在六个物体类别上的平均长时域粒子误差降低了22.8%。该框架进一步支持跨动作、物体实例和场景的迁移,包括从机器人交互到人类操作的零样本迁移,无需目标域动力学拟合。
英文摘要
Deformable-object manipulation is essential for robotic tasks such as folding laundry and handling food, where robots must control shape changes as well as object motion. Predictive soft-body simulation supports these tasks by anticipating deformation under external interactions. However, spatial neighborhoods can misrepresent deformation dependencies, introducing local errors that accumulate over successive predictions. Models fitted to individual scenes must also accommodate changes in object geometry and manipulation conditions. In this work, we propose AIM, an Adaptive Interaction Modeling framework that treats real-to-sim soft-body simulation as a local-global interaction modeling problem. AIM uses motion history and geometry to adapt particle relations over current spatial neighbors and retained connections, while geometry-conditioned global communication coordinates object-wide responses. A unified kinematic control-point interface represents different manipulation configurations, and multi-step autoregressive supervision trains the model on its own predicted trajectories. Experiments on PhysTwin and PGND demonstrate improved motion accuracy and visual fidelity, with a 20.0% reduction in future-prediction tracking error relative to PhysTwin and a 22.8% reduction in mean long-horizon particle error across six object categories relative to PGND. The framework further supports transfer across actions, object instances, and scenes, including zero-shot transfer from robot interactions to human manipulation without target-domain dynamics fitting.