DiagGen:基于仿真诊断的可变形资产智能体生成方法,用于机器人仿真
DiagGen: Agentic Generation of Deformable Assets with Sim-based Diagnostics for Robotic Simulation
- University of British Columbia(不列颠哥伦比亚大学)
- National University of Singapore(新加坡国立大学)
- University of Cambridge(剑桥大学)
- University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
DiagGen通过生成-仿真-诊断-细化循环,利用视觉语言模型智能体在物理仿真器中探测并修复可变形资产,提升仿真就绪性,可直接用于接触丰富的拾取放置任务规划。
AI中文摘要:
虽然适用于仿真的可变形资产对于硅内机器人操作任务至关重要,但现有的生成框架通常在生成之后评估物理合理性,而将物体的仿真响应留作未使用的反馈,无法用于修复上游错误。我们提出了DiagGen,一个智能体框架,通过生成-仿真-诊断-细化循环,将单个野外图像转化为适用于仿真的可变形资产。DiagGen构建部件感知的几何和材料参数,然后使用基于视觉语言模型(VLM)的智能体选择语义信息丰富的区域,在物理仿真器中对其进行探测,观察材料响应,并将有证据支持的修复线索路由到负责的生成阶段。在40个资产上的实验表明,诊断提供了有用的修复线索,并能适度提高生成的可变形资产的质量。最后,我们表明,与可能无法仿真的视觉基础模型生成的资产不同,DiagGen生成的可变形资产可以直接放入高保真物理仿真器中,用于接触丰富的拾取和放置任务的规划与仿真。项目网站为:此HTTPS URL。
英文摘要:
While simulation-ready deformable assets are essential for in-silico robotic manipulation tasks, existing generation frameworks typically assess physical plausibility after generation, leaving an object's simulated response unused as feedback for repairing upstream errors. We present DiagGen, an agentic framework that turns a single in-the-wild image into a simulation-ready deformable asset through a generate--simulate--diagnose--refine loop. DiagGen constructs part-aware geometry and material parameters, then uses a VLM (vision-language model)-based agent to select semantically informative regions, probe them in a physics simulator, observe material responses, and route evidence-backed repair cues to the responsible generation stage. Experiments on 40 assets show that diagnostics provides useful repair cues and can moderately improve the quality of generated deformable assets. Finally, we show that unlike assets generated from visual foundation models which may not be simulatable, DiagGen-generated deformables can be directly dropped into a high-fidelity physical simulator for the planning and simulation of contact-rich pick-and-place tasks. The project's website is https://diaggen.github.io/.