arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于机器人动作表示的语义锚定

Semantic Anchoring for Robotic Action Representations

Yuan Xu, Youheng Shi, Chengyang Li, Wentao Zhu, Yizhou Wang

arXiv 2607.13597首次发表:更新:

发表机构

Peking University; Eastern Institute of Technology, Ningbo; Shanghai Jiao Tong University(北京大学; 宁波东方理工大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究VLA模型微调后动作表示结构受损问题,受镜像神经元理论启发,通过系统探测证实结构变化与任务表现相关。提出即插即用方法,将动作表示锚定到语义流形并分解通道,经多基准测试验证,有效提升了模型在真实世界任务中的表现。

AI 中文摘要

视觉-语言-动作(VLA)模型从预训练的视觉-语言模型继承了丰富的语义表示,但在有限的机器人演示上进行微调会破坏这种结构并削弱泛化能力。由此产生一个基本问题:什么构成良好的动作表示?受镜像神经元理论启发,研究机器人动作表示是否保留预训练编码器捕获的语义结构。系统探测证实微调期间该结构会受损,其质量与任务成功和分布外泛化同步。还引入一种即插即用方法,将动作表示锚定到语义流形,同时将表示分解为共享语义通道和私有通道,推理时丢弃这些通道,使部署模型不变。在不同VLA主干上通过模拟和真实世界基准测试验证,该方法在真实世界分布内任务上提高了18.7%,在分布外泛化上提高了21.5%。

英文摘要

Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limited robot demonstrations degrades this structure and undermines generalization. A fundamental question therefore arises: what constitutes a good action representation? Inspired by the mirror neuron theory's insight that observation and execution share an intention-level encoding, we examine whether a robot's action representations preserve the semantic structure captured by pretrained encoders. Systematic probing confirms that this structure erodes during finetuning, and that its quality synchronizes with both task success and out-of-distribution generalization. We further introduce a plug-and-play method that anchors action representations to a semantic manifold while decomposing representations into a shared semantic channel and a private channel, all discarded at inference, leaving the deployed model unchanged. Validated on different VLA backbones across simulation and real-world benchmarks, our method yields up to +18.7% on real-world in-distribution tasks and +21.5% on out-of-distribution generalization.

CommentsProject Page: https://xy02-05.github.io/SemanticMN

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑