arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35318cs.RO

DexAgent:具有自进化工具库的灵巧操作智能体Human2Sim2Robot框架

DexAgent: An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool Library

  • Stanford University(斯坦福大学)
  • Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Youhui Wang, Yunzhu Li, Li Fei-Fei, Jiajun Wu, Huang Huang

AI总结:

DexAgent是一个智能体化的Human2Sim2Robot框架,通过语义理解、仿真重建、轨迹优化和数据生成四阶段,将单个人类视频转换为机器人训练轨迹,利用自进化工具库和验证器适应多样物体,成功率比基线高3.5倍。

AI中文摘要:

人类视频为灵巧机器人操作提供了可扩展的演示来源。然而,现有的人-仿真-机器人(Human2Sim2Robot)流程依赖预定义的程序,难以适应多样化的物体属性和交互,尤其是涉及铰接和可变形物体的场景。我们提出DexAgent,一个智能体化的Human2Sim2Robot框架,它将单个第一人称视角人类视频和任务提示转换为用于策略训练的物理可行的机器人轨迹。它通过四个阶段运作:人类视频的语义理解、基于属性的仿真重建、机器人轨迹优化和机器人数据生成。在每个阶段,DexAgent通过从其工具库中选择合适的技能或在需要时开发新技能来适应任务和物体属性。特定属性的验证器评估各阶段输出的物理有效性和任务特定要求,并提供反馈以供改进,防止错误在工作流程中传播。这种自适应、验证引导的过程使DexAgent能够处理多样化的物体和长时域任务。在最后阶段,DexAgent在仿真中变化物体和机器人状态,从单个人类视频生成多样化的机器人轨迹,然后对渲染的观测进行重新纹理化以促进仿真到现实的迁移。新开发的技能和验证器保留在其工具库中,使其能够自我进化以积累可复用的能力。随着DexAgent处理更多人类视频,这减少了处理时间。在十一个现实世界任务中,使用DexAgent生成数据训练的策略比竞争基线实现了3.5倍更高的成功率。项目网站:此https URL。

英文摘要:

Human videos offer a scalable source of demonstrations for dexterous robot manipulation. However, existing human-to-simulation-to-robot (Human2Sim2Robot) pipelines rely on predefined procedures that struggle to accommodate diverse object properties and interactions, particularly those involving articulated and deformable objects. We introduce DexAgent, an agentic Human2Sim2Robot framework that converts a single egocentric human video and a task prompt into physically grounded robot trajectories for policy training. It operates through four stages: semantic understanding of human videos, property-based simulation reconstruction, robot trajectory optimization, and robot data generation. At each stage, DexAgent adapts its approach to the task and object properties by selecting suitable skills from its tool library or developing new ones when needed. Property-specific verifiers assess stage outcomes for physical validity and task-specific requirements and provide feedback for refinement, preventing error propagation through the workflow. This adaptive, verification-guided process allows DexAgent to process diverse objects and long-horizon tasks. In the final stage, DexAgent varies object and robot states in simulation to generate diverse robot trajectories from a single human video, then retextures the rendered observations to facilitate sim-to-real transfer. Newly developed skills and verifiers are retained in its tool library, making it self-evolving to accumulate reusable capabilities. This reduces processing time as DexAgent encounters more human videos. Across eleven real-world tasks, policies trained with DexAgent-generated data achieve a 3.5x higher success rate than competing baselines. Project website: https://dexagent1.github.io/.

补充信息

↑