arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

记忆、适应、忽略:训练数据变化下机器人学习机制的诊断

Memorize, Adapt, Ignore: Diagnosing Robot Learning Mechanisms under Training Data Variation

Ke Zhang, Danica J. Sutherland, Chao Liu

arXiv 2609.38401首次发表:更新:

发表机构

PRIME Robotics Lab, Department of Mechanical Engineering, The University of British Columbia; Department of Computer Science, University of British Columbia; Alberta Machine Intelligence Institute(不列颠哥伦比亚大学机械工程系PRIME机器人实验室; 不列颠哥伦比亚大学计算机科学系; 阿尔伯塔机器智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过神经正切核诊断工具,揭示训练数据变化如何使机器人策略从记忆转向适应或忽略,为域随机化设计和模型选择提供实用指导,并经真实硬件实验验证。

AI 中文摘要

训练数据的变化,无论是通过在仿真中设计域随机化(DR)方案,还是为模仿学习整理示范数据,都是提高机器人操作策略鲁棒性的主要手段。然而,其潜在机制仍鲜为人知,从业者通常通过昂贵的试错过程来选择随机化参数。我们通过一系列案例研究来探究这些机制,在ManiSkill中的抓取-放置强化学习以及LIBERO和RoboTwin上对视觉-语言-动作(VLA)模型进行微调等设置中,随机化物体大小、颜色和类型以及场景光照和语言提示。我们使用经验神经正切核(NTK)作为主要诊断工具,检查模型行为和内部表示。我们表明,NTK能够区分内部学习机制从“记忆”不同情况但变化不足(例如,学习对大立方体做什么,以及对小立方体做什么)到“适应”当前情况且变化充分的转变。基于NTK的信噪比也有助于区分策略何时学会了“忽略”与任务无关的因素(例如,将蓝色和红色立方体等同对待,而不是学习一个蓝色子策略和一个红色子策略)。我们利用这些诊断方法,为设计域随机化方案、选择模型和检测捷径学习提供了实用指导。我们进一步比较了不同类型的表示,并使用基于ACT的模仿学习在真实世界硬件实验中验证了我们的发现。

英文摘要

Training data variation, whether through designing a domain randomization (DR) scheme in simulation or curating demonstrations for imitation learning, is a primary lever for improving the robustness of robotic manipulation policies. Yet its underlying mechanisms remain poorly understood, and practitioners typically select randomization parameters through expensive trial and error. We investigate these mechanisms through a series of case studies, randomizing object size, color, and type as well as scene lighting and linguistic prompts across settings including pick-and-place RL in ManiSkill and fine-tuning of vision-language-action (VLA) models on LIBERO and RoboTwin. We examine both model behavior and internal representations, using the empirical neural tangent kernel (NTK) as our primary diagnostic tool. We show that the NTK distinguishes a shift in the internal learning mechanism from \textit{memorizing} different situations with insufficient variation (e.g.\ learning what to do for a large cube, and what to do for a small cube) to \textit{adapting} to the situation at hand with sufficient variation. An NTK-based signal-to-noise ratio also helps distinguish when policies have learned to \emph{ignore} task-irrelevant factors (e.g.\ treating blue and red cubes identically, instead of learning a blue sub-policy and a red sub-policy). We use these diagnostics to develop practical guidance for designing DR schemes, selecting models, and detecting shortcut learning. We further compare different kinds of representations and validate our findings with real-world hardware experiments using ACT-based imitation learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑