arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向视觉-语言-动作模型的鲁棒指令泛化的基础语义重绑定

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models

Zhaokai Yin, Zhipeng Zhang

arXiv 2608.02497首次发表:更新:

发表机构

School of Artificial Intelligence, Shanghai Jiao Tong University; Anyverse Dynamics(上海交通大学人工智能学院; Anyverse Dynamics公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉-语言-动作模型因指令释义导致性能骤降的架构问题,提出GSR方法,在LIBERO-Para基准提升成功率44.6%,推出ParaVLA模型,规避低效数据扩展,实现鲁棒指令泛化。

AI 中文摘要

视觉-语言-动作(VLA)模型在机器人操纵任务中表现出色,但当规范指令被简单释义时,其性能会急剧下降。尽管这种脆弱性通常通过代价高昂的数据扩展来解决,但我们的探究表明,其根本原因是架构层面的问题,而非语义理解不足。具体而言,我们发现当前的VLA模型在内部能够成功保留正确的任务身份,其失效实际上源于动态视觉观测与文本的联合编码,这会引入系统性的特征偏移。由于下游动作策略对这些变化高度敏感,它无法将保留的语义转化为正确的控制指令。为解决这一结构性瓶颈,我们提出了基础语义重绑定(Grounded Semantic Re-binding, GSR),这是一种巧妙的干预方法,通过显式融合独立提取的任务语义与原生视觉特征,从头开始训练完全重新初始化的动作专家,从而绕过不稳定的联合路由。这种针对性干预仅使用规范演示就能显著恢复释义不变性。在LIBERO-Para基准测试中,GSR将成功率提升了多达44.6%,使轻量级模型能够与大规模基线模型相媲美,并将最先进模型的PRIDE分数推向70.4的新纪录,在指令泛化能力上优于最近推出的大规模预训练模型Xiaomi-Robotics-0。基于这些发现,我们还引入了ParaVLA,这是一个原生解耦的0.33B参数模型,对指令重述表现出近乎完美的鲁棒性。最终,我们的工作证明,通过巧妙的结构设计即可实现鲁棒的语义基础,从而规避低效的蛮力数据扩展范式。

英文摘要

Vision-Language-Action (VLA) models excel in robotic manipulation but suffer catastrophic performance drops when canonical instructions are simply paraphrased. Although this brittleness is typically addressed through costly data scaling, our probing reveals that the root cause is architectural rather than a lack of semantic understanding. Specifically, we demonstrate that current VLAs successfully retain the correct task identity internally. The failure actually stems from the joint encoding of dynamic visual observations and text, which introduces systematic feature shifts. Because the downstream action policy is highly vulnerable to these variations, it fails to translate the preserved semantics into correct control commands. To resolve this structural bottleneck, we propose Grounded Semantic Re-binding (GSR), an elegant intervention that bypasses unstable joint routing by explicitly fusing independently extracted task semantics with native visual features to train a completely re-initialized action expert from scratch. This targeted intervention dramatically restores paraphrastic invariance using only canonical demonstrations. On the LIBERO-Para benchmark, GSR improves success rates by up to 44.6 percent. It enables lightweight models to rival massively scaled baselines and pushes state-of-the-art models to a new record PRIDE score of 70.4, outperforming the recently introduced large-scale pretrained model Xiaomi-Robotics-0 in instruction generation capabilities. Building on these insights, we also introduce ParaVLA, a natively decoupled 0.33B-parameter model exhibiting near-perfect robustness to instruction rewording. Ultimately, our work proves that robust semantic grounding can be achieved through elegant structural design, bypassing the inefficient brute-force data scaling paradigm.

Comments23 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑