arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EigenDEXplore:基于人类先验的灵巧操作结构化探索

EigenDEXplore: Structured Exploration for Dexterous Manipulation with Human Priors

Harsh Gupta, Tyler Ga Wei Lum, Changhao Wang, Chuer Pan, C. Karen Liu, Jeannette Bohg, Shuran Song

arXiv 2610.07681首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对灵巧操作的高维探索难题,提出EigenDEXplore方法,利用人类运动先验沿特征向量添加相关扰动,在不改变动作空间的情况下提升探索效率,在多种任务和设置中优于现有基线。

AI 中文摘要

灵巧操作提出了一个具有挑战性的高维优化问题,因为有用的行为需要多个手部关节的协调运动。在强化学习(RL)和基于采样的轨迹优化中,探索通常依赖于独立的机器人关节扰动,这使得协调行为难以被发现。先前的工作利用从人类手部数据中学习的协调关节运动的低维空间来减少抓取学习的搜索空间,但这限制了通用操作所需的表达能力。一些方法将学习到的动作与关节空间动作相结合以恢复表达能力,但这增加了维度并引入了冗余。我们在多种操作设置中研究了这些影响,变化了动作维度、探索策略和人类数据的来源。我们的实验表明,人类运动先验在用于结构化探索而非改变动作表示时最为有效。基于这一发现,我们提出了EigenDEXplore,它通过向独立的关节空间噪声中添加沿人类特征向量的扰动来诱导相关探索,同时保持动作空间不变。在多个灵巧手上,EigenDEXplore在抓取、手内重定向和接触丰富的操作中始终优于关节空间和学习到的动作空间基线。这些优势涵盖了无结构和参考引导的RL、轨迹优化以及仿真到现实的部署,并且在奖励塑造和课程设计较少的环境中最为显著。

英文摘要

Dexterous manipulation poses a challenging high-dimensional optimization problem, as useful behaviors require coordinated motion across many hand joints. In reinforcement learning (RL) and sampling-based trajectory optimization, exploration commonly relies on independent robot joint perturbations, making coordinated behaviors difficult to discover. Prior work reduces this search space for grasp learning using low-dimensional spaces of coordinated joint motions learned from human hand data, but this restricts the expressivity required for general manipulation. Some combine learned and joint-space actions to restore expressivity, but this increases dimensionality and introduces redundancy. We study these effects across diverse manipulation settings, varying action dimensionality, exploration strategy, and the source of human data. Our experiments suggest that human-motion priors are most effective when used to structure exploration rather than change the action representation. Motivated by this finding, we propose EigenDEXplore, which induces correlated exploration by adding perturbations along human-derived eigenvectors to independent joint-space noise, leaving the action space unchanged. Across multiple dexterous hands, EigenDEXplore consistently outperforms joint-space and learned action-space baselines in grasping, in-hand reorientation, and contact-rich manipulation. These gains span unstructured and reference-guided RL, trajectory optimization, and sim-to-real deployment, and are largest in settings with less reward shaping and curriculum design.

Comments15 pages, 12 figures. Project page: https://eigendexplore.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑