arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AffordCraft:从单张图像可扩展构建任务就绪的仿真资产

AffordCraft: Scalable Construction of Task-Ready Simulation Assets from Single Images

Haoyun Yang, Xueyang Zhou, Ziyi Xie, Yongchao Chen

arXiv 2610.06643首次发表:更新:

发表机构

College of AI, Tsinghua University; School of Computer Science, Chongqing University; School of Cyber Science and Engineering, Huazhong University of Science and Technology; College of Design and Engineering, National University of Singapore(清华大学人工智能学院; 重庆大学计算机学院; 华中科技大学网络空间安全学院; 新加坡国立大学设计与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AffordCraft通过检索而非生成,从单张图像和任务指令构建物理有效的仿真资产,显著提高成功率并降低计算成本,支持可扩展库和下游操作任务。

AI 中文摘要

仿真中的机器人学习依赖于仿真器提供的物体。许多任务需要具有独立部件、允许所需运动的关节以及接触下保持有效的物理属性的物体。现有方法对每张图像重新恢复这种结构:生成模型预测的部件和关节大多在仿真中无法稳定或移动,而通用智能体对每张照片需要长时间的模型调用会话。AffordCraft通过检索而非生成,从单张RGB图像和任务指令构建此类资产:它定位要操作的物体和部件,从铰接资产库中选择匹配条目,并在保持其部件和关节完整的同时将其适配到图像。在没有任何框或掩码标记物体的情况下,AffordCraft为来自31个类别的2000张照片中的1703张生成了物理有效的资产。五种生成方法在同一批照片上最多通过45%,并且在有效资产的中位GPU时间上是我们的10到78倍。在50张杂乱图像上,自动检测后,237个标注物体中有162个通过了相同的物理测试。将库从141个条目扩展到11372个条目无需改变方法,并将类别覆盖率从46%提高到100%,具有请求标签的选择比例从18%提高到51%。我们还从构建的资产中构建操作任务,包括单个物体和组合场景;基于脚本演示训练的策略能够从未在训练中出现的初始状态完成这两种任务。

英文摘要

Robot learning in simulation depends on the objects the simulator offers. Many tasks need objects with separate parts, joints that allow the required motion, and physical properties that remain valid under contact. Existing methods recover this structure anew for every image: generative models predict parts and joints that mostly fail to settle or move in simulation, and general-purpose agents need a long session of model calls for each photograph. AffordCraft builds such an asset from a single RGB image and a task instruction by retrieval instead of generation: it locates the object and the part to operate, selects a matching entry from a library of articulated assets, and fits it to the image while keeping its parts and joints intact. Without any box or mask marking the object, AffordCraft produces a physically valid asset for 1,703 of 2,000 photographs from 31 categories. Five generative methods pass on at most 45% of the same photographs and, at the median, need 10 to 78 times our GPU time per valid asset. On 50 cluttered images, 162 of 237 annotated objects pass the same physical test after automatic detection. Growing the library from 141 to 11,372 entries needs no change to the method and raises category coverage from 46% to 100% and the share of selections with the requested label from 18% to 51%. We also build manipulation tasks from the constructed assets, both with single objects and in composed scenes; policies trained on scripted demonstrations complete both kinds of tasks from initial states unseen in training.

Comments33 pages, 14 figures, 17 tables. Project page: https://affordcraft.github.io Code: https://github.com/AffordCraft/AffordCraft

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑