arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Agentic RSR:通过场景重建和执行接地机器人策略实现真实到仿真到真实

Agentic RSR: Real-to-Sim-to-Real through Scene Reconstruction and Execution-Grounded Robot Policies

Yihan Li, Yating Feng, Shengjiu Sun, Jianing Chen, Hao Ren, Bowen Yang, Weisheng Xu, Qiwei Wu, Hui Cheng, Renjing Xu

arXiv 2610.10479首次发表:更新:

发表机构

Sun Yat-sen University; The Hong Kong University of Science and Technology (Guangzhou)(中山大学; 香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Agentic RSR框架通过场景重建与执行接地策略,实现真实到仿真到真实的闭环,在18个场景中深度误差0.1057米,真实机器人任务成功率保留仿真性能的80%。

AI 中文摘要

真实机器人工作空间的仿真必须保留与任务相关的交互,而在仿真中开发的策略必须能处理真实机器人可获得的观测。然而,场景重建和策略开发往往被分开处理。我们提出了Agentic Real-to-Sim-to-Real(Agentic RSR),一个通过同一操作任务将场景重建、策略开发和真实机器人执行联系起来的框架。给定工作空间视频、任务描述和已知的机器人模型,一个智能体恢复度量尺度,使用视觉反馈迭代细化场景,并在MuJoCo中检查与任务相关的交互。然后,一个编码智能体开发一个可执行的策略,从特权物体姿态逐步过渡到视觉观测和随机化仿真。该策略可以在一次调用中交错多个观测和动作,而智能体使用执行反馈来继续、重试或修改其方法。一个共享的任务级接口将策略和积累的经验带到真实机器人,在那里新的观测和安全检查指导执行。在涉及两个机器人的18个重建场景中,平均四视图深度MAE相对于参考深度估计为0.1057米,平均Lab ΔE76为11.04,平均灰度SSIM为0.6990。在真实机器人实验中,总体任务成功率达到了仿真任务成功率的80%,表明硬件上保留了大量的仿真性能。代码和重建场景数据将公开提供。

英文摘要

A simulation of a real robot workspace must preserve task-relevant interactions, while policies developed in it must operate on observations available to the real robot. Yet scene reconstruction and policy development are often treated separately. We present Agentic Real-to-Sim-to-Real (Agentic RSR), a framework that links scene reconstruction, policy development, and real-robot execution through the same manipulation task. Given a workspace video, a task description, and a known robot model, an agent recovers metric scale, iteratively refines the scene using visual feedback, and checks task-relevant interactions in MuJoCo. A coding agent then develops an executable policy, progressing from privileged object poses to visual observations and randomized simulation. The policy can interleave multiple observations and actions within one invocation, while the agent uses execution feedback to continue, retry, or revise its approach. A shared task-level interface carries the policy and accumulated experience to the real robot, where fresh observations and safety checks guide execution. Across 18 reconstructed scenes involving two robots, the mean four-view Depth MAE against reference depth estimates is 0.1057 m, the mean Lab $ΔE_{76}$ is 11.04, and the mean grayscale SSIM is 0.6990. In real-robot experiments, the aggregate task success rate reaches 80% of the simulation task success rate, indicating substantial retention of simulated performance on hardware. Code and reconstructed scene data will be made publicly available.

Comments25 pages including appendices, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑