arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

如何从人类会避免的行为中学习?用于灵巧操作的真实世界强化学习中的干预感知世界模型

How to Learn from What a Human Would Avoid? Intervention-Aware World Models with Real-World RL for Dexterous Manipulation

Jiaju Yin, Zhenhui Zhang, Lixin Xu, Heng Zhang, Jun Shao, Yating Feng, Arash Ajoudani, Renjing Xu

arXiv 2609.06009首次发表:更新:

发表机构

HKUST (Guangzhou); Italian Institute of Technology; Zhejiang University(香港科技大学(广州); 意大利理工学院; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对真实世界灵巧操作中硬件故障代价高昂的问题,提出干预感知世界模型WHIRL,将人类干预转化为预测信号,通过风险塑形使16自由度LEAP手在复杂抓取任务中达到96.7%成功率,并减少84%干预负担。

AI 中文摘要

多指灵巧操作由于高维动作空间和硬件故障的昂贵代价,仍然是真实世界强化学习(RL)的前沿挑战。虽然人在回路(HIL)强化学习允许操作员在故障发生前进行干预,但当前的流程通常将这些干预视为反应性修正,丢弃了操作员决定接管控制所固有的丰富安全信号。在本文中,我们提出疑问:我们如何从人类会避免的行为中学习?我们提出了WHIRL,一个安全感知的强化学习框架,将二元人类干预转化为前向预测信号,用于主动风险规避。我们的方法核心是一个干预感知的潜在世界模型,具有四个预测头:动力学、奖励、终止,以及一个新颖的逐状态干预概率头,该头学习预测未来状态下人类接管的可能性。这个头提供了一个行动者侧的风险塑形项,通过模拟操作员的内部安全阈值,阻止策略进入“易干预”区域。我们在一个16自由度LEAP手上评估了我们的框架,任务涵盖凸形和不规则物体抓取、棱柱形操作以及长时程多阶段任务。我们的结果表明,预测性风险塑形使系统在复杂抓取任务中实现了96.7%的成功率,同时在按步加权项下将操作员干预负担减少了高达84%。通过在人直觉与预测性世界建模之间闭环,这项工作为在真实世界中训练复杂灵巧智能体提供了一种实用的安全感知方法,同时减少了操作员疲劳和硬件风险暴露。

英文摘要

Multi-fingered dexterous manipulation remains a frontier for real-world reinforcement learning (RL) due to the high-dimensional action space and the prohibitive cost of hardware failures. While human-in-the-loop (HIL) RL allows operators to intervene before failures occur, current pipelines often treat these interventions as reactive corrections, discarding the rich safety signal inherent in the operator's decision to take control. In this paper, we ask: How can we learn from what a human would avoid? We present WHIRL, a safety-aware RL framework that transforms binary human interventions into forward-predictive signals for proactive risk avoidance. Our approach centers on an intervention-aware latent world model with four prediction heads: dynamics, reward, termination, and a novel per-state intervention-probability head that learns to predict the likelihood of a human takeover at future states. This head provides an actor-side risk-shaping term that discourages the policy from entering "intervention-prone" regions, modeling the operator's internal safety threshold. We evaluate our framework on a 16-DoF LEAP Hand across tasks spanning convex and irregular object grasping, prismatic manipulation, and long-horizon multi-stage tasks. Our results show that predictive risk-shaping enables the system to achieve a 96.7 percent success rate on complex grasping tasks while reducing the operator intervention burden by up to 84 percent in step-weighted terms. By closing the loop between human intuition and predictive world modeling, this work provides a practical safety-aware recipe for training complex dexterous agents in the real world while reducing operator fatigue and hardware-risk exposure.

Comments15 pages, 8 figures, 3 tables. Accepted by CoRL 2026. Project page: https://whirl-dexterous.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑