发表机构
AXXX; MIRAI; MISIS(AXXX; 未来人工智能研究机构; 俄罗斯国立科技大学莫斯科钢铁合金学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究能否通过优化语言空间而非更新动作权重来改进冻结的 VLA 策略,提出语言条件空间策略,用强化学习优化,实验表明该优化能提升特定任务成功率,证明语言可作为机器人基础模型的可优化变量。
AI 中文摘要
视觉 - 语言 - 动作(VLA)模型通常被视为基于自然语言任务描述的端到端动作策略。然而,其行为常强烈依赖指令表述方式,表明语言不仅是任务标签,还是可优化的条件输入。我们研究能否通过优化语言空间而非更新动作权重来改进冻结的 VLA 策略。我们的方法引入了一种语言条件空间策略,该策略使用对象外观、空间关系和目标接地线索将人类指令转换为简短的 VLA 接地命令。语言条件空间策略以失败派生的命令空间先验初始化,并通过从稀疏任务完成奖励中进行强化学习进行优化,而下游 VLA 保持完全冻结。这产生了语言条件空间优化:强化学习发现哪些 VLA 接地命令能从冻结的动作策略中最好地引发成功行为。在 RL4VLA 和 VL-Think 上的实验表明,语言条件空间优化提高了对指令敏感、符号和多对象操纵任务的成功率,证明语言可以作为机器人基础模型的可优化变量。网站:this https URL
英文摘要
Vision-Language-Action (VLA) models are commonly treated as end-to-end action policies conditioned on natural-language task descriptions. In practice, however, their behavior often depends sharply on how the instruction is phrased, suggesting that language is not merely a task label but an optimizable conditioning input. We study whether frozen VLA policies can be improved by optimizing language space rather than updating action weights. Our method introduces a language-conditioning space policy that translates a human instruction into a short VLA-grounded command using object appearance, spatial relations, and target-grounding cues. The language-conditioning space policy is optimized with reinforcement learning from sparse task-completion rewards, while the downstream VLA remains fully frozen. Experiments on RL4VLA and VL-Think show that language-conditioning space optimization improves success on instruction-sensitive, symbolic, and multi-object manipulation tasks, demonstrating that language can serve as an optimizable variable for robot foundation models. Website: https://cognitiveaisystems.github.io/VLA-Grounder/