战略扩展:通过偏差感知评估和数据收集学习机器人操作中的组合泛化
Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation
浏览论文内容
中文总结 AI 辅助
研究机器人操作中组合泛化问题,通过引入诊断框架量化指令因素偏差,提出偏差感知数据收集策略,能定位预训练策略捷径问题,实现更高效可泛化的策略学习。
中文摘要 AI 辅助
组合泛化对机器人遵循各种指令至关重要。然而,预训练策略会走捷径,依赖显著线索而非基于语言。我们引入一个诊断框架,将这种失败定位到各个“指令因素”,如颜色、动词、物体、大小和空间属性等可重复使用的语义组件。该框架形式化了指令因素偏差,即微调策略过度依赖主导因素走捷径的倾向,并通过两个指标量化:因素主导率(FDR)捕捉因素间的成对偏差,因素主导层次结构(FDH)将其汇总为全局排名。对六个基础策略的评估揭示了大致一致的排序,即颜色≥物体≥空间≥动词≥大小,颜色占主导,动词和大小最缺乏基础。我们进一步表明这种诊断是可行的:一种偏差感知数据收集策略,将固定预算重新分配到缺乏基础的因素上,在模拟和真实机器人上使用一半演示时优于基线,从而实现更具样本效率和可泛化的策略学习。
英文摘要
Compositional generalization is essential for robot to follow diverse instructions. However, pretrained policies are known to take shortcuts, deferring to salient cues rather than grounding language. We introduce a diagnostic framework that localizes this failure to individual \textit{instruction factors}, \textit{e.g.,} reusable semantic components such as color, verb, object, size, and spatial attribute. Our framework formalizes instruction factor bias, the tendency of fine-tuned policies to over-rely on dominant factors as shortcuts, and quantifies it through two metrics: Factor Dominance Rate (FDR), capturing pairwise bias between factors, and Factor Dominance Hierarchy (FDH), aggregating these into a global ranking. Evaluation on six foundation policies reveals broadly consistent ordering, \textit{i.e.}, color $\geq$ object $\geq$ spatial $\geq$ verb $\geq$ size, with color dominant, and verb and size most under-grounded. We further show the diagnosis is actionable: a bias-aware data collection strategy that reallocates a fixed budget toward under-grounded factors outperforms baselines in simulation and on a real robot using half the demonstrations, thereby enabling more sample-efficient and generalizable policy learning.
发表机构
- Northeastern University(东北大学)
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。