arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38225cs.ROcs.AI

SynIL:利用协同性从不完美演示数据集进行离线模仿学习

SynIL: Leveraging Synergy for Offline Imitation Learning from Imperfect Demonstration Datasets

Yuto Tanaka, Kyo Kutsuzawa, Martina Doku, Dai Owaki, Mitsuhiro Hayashibe

首次发表
浏览论文内容

中文总结 AI 辅助

SynIL提出基于协同性的无标签演示质量评估框架,通过量化运动协同性生成密集奖励信号,在D4RL和Robomimic基准上超越行为克隆,媲美甚至优于真实奖励训练的离线强化学习。

中文摘要 AI 辅助

模仿学习使机器人能够直接从大规模演示数据集中获取复杂技能,但当数据集被次优或噪声演示污染时,其性能会严重下降。虽然先前的质量评估方法尝试过滤或重新加权数据,但它们通常依赖于人工预选专家参考数据或任务特定启发式规则,限制了可扩展性。为应对这一挑战,我们提出了SynIL(基于协同性的模仿学习),一种用于离线强化学习中自动化、无标签演示质量评估的新颖框架。基于神经科学证据——运动协同性(一种运动中的低维协调结构)与运动熟练度直接相关——SynIL通过算法量化协同性表现,利用自监督奖励回归生成密集的、转换级别的奖励信号。在D4RL运动基准和多人类Robomimic操作数据集上的全面评估表明,协同性派生的奖励与真实奖励强相关。此外,SynIL显著优于行为克隆(BC),并且在稀疏奖励的人类遥操作场景中,其性能可与基于真实环境奖励训练的离线强化学习相媲美,甚至更优。

英文摘要

Imitation learning enables robots to acquire complex skills directly from massive demonstration datasets, but its performance degrades severely when datasets are contaminated with suboptimal or noisy demonstrations. While prior quality-assessment methods attempt to filter or reweight data, they typically rely on manual pre-selection of expert reference data or task-specific heuristics, limiting scalability. To address this challenge, we introduce SynIL (Synergy-based Imitation Learning), a novel framework for automated, label-free demonstration quality assessment in offline reinforcement learning. Grounded in neuroscientific evidence that motor synergy, a low-dimensional coordinated structure in movement, correlates directly with motor proficiency, SynIL algorithmically quantifies synergy manifestation to generate dense, transition-level reward signals via self-supervised reward regression. Comprehensive evaluations on D4RL locomotion benchmarks and multi-human Robomimic manipulation datasets demonstrate that synergy-derived rewards correlate strongly with ground-truth rewards. Furthermore, SynIL substantially outperforms Behavior Cloning (BC) and achieves performance comparable to, and in sparse-reward human teleoperation scenarios, superior to, offline reinforcement learning trained on true environment rewards.

发表机构

  • Tohoku University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

↑