arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体真实到模拟:使用视觉语言智能体进行基于物理的世界建模

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

Guanxiong Chen, Qianjun Xia, Jiawei Peng, Heng Zhang, Pengyu Jing, Bole Ma, Justin Qian, Yixian Cheng, Ziyi Jiao, Bingyang Zhou, Yiduo Qu, Luoxin Ye, Kaifeng Zhang, Kunyi Wang, Weijia Zeng, Yunuo Chen, Pengzhi Yang, Ziqiu Zeng, Siyuan Luo, Huamin Wang, Chao Liu, Alan Yuille, Fan Shi, Changxi Zheng, Yunzhu Li, Chenfanfu Jiang, Peter Yichen Chen

arXiv 2607.19190首次发表:更新:

发表机构

University of British Columbia; Johns Hopkins University; National University of Singapore; Columbia University; University of California, Los Angeles; Style3D(英属哥伦比亚大学; 约翰·霍普金斯大学; 新加坡国立大学; 哥伦比亚大学; 加利福尼亚大学洛杉矶分校; 无合适对应中文名)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对机器人与物体交互的真实到模拟转换难题,提出智能体真实到模拟框架,利用视觉语言智能体进行广义物理世界建模,能转换真实记录为可模拟孪生体,在多场景评估效果良好,成本低,可用于下游机器人任务。

AI 中文摘要

机器人与物体交互的真实到模拟转换仍然劳动密集,因为它不仅需要视觉重建。一个简化的真实到模拟过程必须恢复场景几何形状和物体状态,推断物理参数,并将参与者、物体、相机、姿态和轨迹组装成可运行的物理模拟。目前该过程仍依赖于视觉基础模型的手动调整、网格清理、坐标框架对齐以及跨视觉感知工具和模拟器的脆弱工作流程粘合。我们引入了智能体真实到模拟,这是一个使用视觉语言智能体进行广义物理世界建模的框架,将物体 - 机器人交互的真实世界记录转换为可模拟的情节孪生体,保留观察结果、几何形状、机器人交互和物体状态。我们在刚体操作、可变形物体交互和人形运动场景上评估了智能体真实到模拟,跨越了通常由单独的真实到模拟管道处理的领域,朝着可扩展转换迈出了第一步。该框架的智能体决策可以由开放权重的VLM后端驱动,成本仅为前沿模型的一小部分,同时实现相当的转换成功率。我们旨在将生成的与真实世界对齐的孪生体用于下游机器人任务,特别是策略学习和评估。项目网站可在这个https网址获取。

英文摘要

Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters, and assemble actors, objects, cameras, poses, and trajectories into a runnable physical simulation. Today this process still depends on brittle workflow glue across visual perception tools and simulators: manual tuning of visual foundation models, mesh cleanup, coordinate frame alignments, etc. We introduce \textit{Agentic Real2Sim}, a framework for generalized physical world modeling with vision-language agents that converts a real-world recording of object-robot interaction into a simulatable episodic twin, and connects the resulting twin to downstream policy fine-tuning and evaluation. We evaluate Agentic Real2Sim on rigid-object manipulation, deformable-object interaction, and humanoid motion scenes, spanning domains that are usually handled by separate Real2Sim pipelines. The framework's agentic decisions can be driven by an open-weight VLM backend at a small fraction of the cost of frontier models, while attaining a comparable conversion success rate. The framework further supports custom scene conversion, fine-tuning of a pretrained policy with data generated from converted episodes, and works effectively as a surrogate for real-world policy evaluation. The project site, including code is available at https://agentic-real2sim.github.io.

CommentsPost conf sub update

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑