发表机构
Purdue University; Indiana University Bloomington(普渡大学; 印第安纳大学伯明顿分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种语言扰动基准,用于评估持续模仿学习中机器人策略的语言根基,发现强持续学习性能并不总能保证语言引导的可靠性。
AI 中文摘要
持续模仿学习评估机器人能否在不遗忘先前所学技能的情况下学习新知识。然而,保持任务性能并不能确保行为仍然扎根于语言,因为策略可能依赖场景线索、物体关联或记忆的任务结构。我们引入一种基准协议,研究当机器人策略学习连续任务时,语言引导行为如何变化。我们为LIBERO的Goal、Spatial、Object和Long套件构建了保持意义和改变意义的指令变体。策略实验聚焦于LIBERO-Goal,在每个持续学习阶段后评估原始指令和释义指令。我们比较了具有代表性的持续模仿学习方法在其原始假设下的表现,同时将任务能力与语言敏感性分开。所提出的诊断方法通过测量语义鲁棒性、目标适应性和语言敏感性,补充了标准的学习和遗忘指标。结果表明,强大的持续学习性能并不总能转化为可靠的语言根基,我们的诊断有助于确定保留的技能是否仍由其指令正确引导。附加材料可在该https URL获取。
英文摘要
Continual imitation learning evaluates whether a robot can learn new knowledge without forgetting previously learned skills. However, retaining task performance does not ensure the behavior remains grounded in language because policies may rely on scene cues, object associations, or memorized task structure. We introduce a benchmark protocol to study how language-guided behavior changes as robotic policies learn successive tasks. We construct meaning-preserving and meaning-changing instruction variants for the Goal, Spatial, Object, and Long suites of LIBERO. Policy experiments focus on LIBERO-Goal, evaluating Original and Paraphrase instructions after each continual-learning stage. We compare representative continual imitation learning methods under their original assumptions while separating task competence from language sensitivity. The proposed diagnostics complement standard learning and forgetting metrics by measuring semantic robustness, goal adaptation, and language sensitivity. Results show that strong continual-learning performance does not always translate to reliable language grounding, and our diagnostics help determine whether retained skills remain correctly guided by their instructions. Additional materials are available at https://sites.google.com/view/stillgrounded