arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.17783cs.RO

RoboSeek:你需要与你的物体进行交互

RoboSeek: You Need to Interact with Your Objects

  • FNii-Shenzhen(FNii-深圳)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Northeastern University(东北大学)
  • Harbin Engineering University(哈尔滨工程大学)
  • Infused Synapse AI

机构由 AI 辅助整理,请以论文原文为准。

Yibo Peng, Jiahao Yang, Shenhao Yan, Ziyu Huang, Shuang Li, Shuguang Cui, Yiming Zhao, Yatong Han

更新

AI总结:

RoboSeek提出一种基于交互经验的具身动作执行框架,通过仿真闭环训练优化视觉先验,并利用real2sim2real迁移实现鲁棒真实世界执行,在八个长时程操作任务上平均成功率达79%,显著超越基线方法。

AI中文摘要:

通过探索和交互来优化与细化动作执行,是机器人操作中一种很有前景的方式。然而,以交互为驱动的机器人学习的实用方法仍未被充分探索,尤其是在长时程任务中,序贯决策、物理约束和感知不确定性带来了重大挑战。受具身认知理论的启发,我们提出了RoboSeek,这是一个用于具身动作执行的框架,利用交互经验来完成操作任务。RoboSeek通过在仿真中进行闭环训练来优化来自高层感知模型的先验知识,并通过real2sim2real迁移流程实现鲁棒的真实世界执行。具体来说,我们首先使用三维重建在仿真中复现真实世界环境,以提供视觉和物理上一致的环境,然后利用视觉先验,在仿真中使用强化学习和交叉熵方法训练策略。学习到的策略随后部署到真实机器人平台上执行。RoboSeek与硬件无关,并在多个机器人平台上针对八个长时程操作任务进行了评估,这些任务涉及序贯交互、工具使用和物体处理。我们的方法实现了79%的平均成功率,显著优于成功率仍低于50%的基线方法,凸显了其跨任务和跨平台的泛化性与鲁棒性。实验结果验证了我们的训练框架在复杂、动态的真实世界环境中的有效性,并证明了所提出的real2sim2real迁移机制的稳定性,为更具泛化能力的具身机器人学习铺平了道路。项目页面:https://russderrick.github.io/Roboseek/

英文摘要:

Optimizing and refining action execution through exploration and interaction is a promising way for robotic manipulation. However, practical approaches to interaction-driven robotic learning are still underexplored, particularly for long-horizon tasks where sequential decision-making, physical constraints, and perceptual uncertainties pose significant challenges. Motivated by embodied cognition theory, we propose RoboSeek, a framework for embodied action execution that leverages interactive experience to accomplish manipulation tasks. RoboSeek optimizes prior knowledge from high-level perception models through closed-loop training in simulation and achieves robust real-world execution via a real2sim2real transfer pipeline. Specifically, we first replicate real-world environments in simulation using 3D reconstruction to provide visually and physically consistent environments, then we train policies in simulation using reinforcement learning and the cross-entropy method leveraging visual priors. The learned policies are subsequently deployed on real robotic platforms for execution. RoboSeek is hardware-agnostic and is evaluated on multiple robotic platforms across eight long-horizon manipulation tasks involving sequential interactions, tool use, and object handling. Our approach achieves an average success rate of 79%, significantly outperforming baselines whose success rates remain below 50%, highlighting its generalization and robustness across tasks and platforms. Experimental results validate the effectiveness of our training framework in complex, dynamic real-world settings and demonstrate the stability of the proposed real2sim2real transfer mechanism, paving the way for more generalizable embodied robotic learning. Project Page: https://russderrick.github.io/Roboseek/

↑