测量机器人策略中的语言迁移:向 Cosmos3 视觉-语言-动作策略添加希腊语
Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy
- Sophea AI(索菲亚人工智能公司)
- KIEFER SA(基弗股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过向Cosmos3策略添加希腊语,发现测量语言迁移的关键在于构建零基线和跨种子复制,双语训练可稳定提升6.7-7.1个百分点。
AI中文摘要:
机器人基础模型主要在英语环境下进行训练和评估,而大多数语言不存在机器人演示语料库。我们研究仅使用机器改写指令且不改变架构的情况下,向一个开放的视觉-语言-动作栈添加希腊语。主要挑战在于测量而非翻译。几种看似合理的工具会产生错误结论:颜色直方图指标奖励噪声,单目标基准在正确希腊语下得分为84.6%,在故意错误的指令下得分为82.6%,训练损失无法预测希腊语的成功,单次运行比较受种子变化主导。在一个包含九十项任务的判别性测试套件中,每臂使用三个种子,没有希腊语演示的多语言文本塔仍处于其错误指令的下限,而仅希腊语训练相对于其对照组最多超出2.7个百分点。双语训练相对于其对照组产生一致的6.7-7.1个百分点的优势,并达到英语性能的大约五分之二。该策略还过度拟合了翻译器的措辞;对每项任务使用七种措辞进行训练可将这一惩罚大约减半。从语言适应的世界模型进行热启动以及解冻文本塔都会降低性能。结果支持低资源机器人策略本地化的两个实际要求:在信任指标之前构建一个保证的零基线,并在多个种子上复制低资源语言的结果。
英文摘要:
Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do not exist for most languages. We study the addition of Greek to an open vision-language-action stack using only machine-rephrased instructions and no architecture changes. The main challenge is measurement rather than translation. Several plausible instruments produce false conclusions: a color-histogram metric rewards noise, a single-goal benchmark scores 84.6% under correct Greek and 82.6% under deliberately wrong instructions, training loss fails to predict Greek success, and single-run comparisons are dominated by seed variation. On a discriminative ninety-task suite with three seeds per arm, a multilingual text tower without Greek demonstrations remains at its wrong-instruction floor, while Greek-only training exceeds its control by at most 2.7 points. Bilingual training yields a consistent 6.7-7.1 point margin over its control and reaches about two fifths of English performance. The policy also overfits the translator's phrasing; training on seven phrasings per task approximately halves this penalty. Warm-starting from a language-adapted world model and unfreezing the text tower both degrade performance. The results support two practical requirements for low-resource robot-policy localization: build a guaranteed null before trusting a metric, and replicate low-resource-language results across seeds.