arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向视觉-语言-动作模型的来自解析概念的显式运动学引导

Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models

Mingyang Sun, Jiude Wei, Xiujian Liang, Qichen He, Donglin Wang, Cewu Lu, Jianhua Sun

arXiv 2607.26513首次发表:更新:

发表机构

Zhejiang University; Shanghai Innovation Institute; Fudan University; Westlake University; Shanghai Jiao Tong University(浙江大学; 上海创新研究院; 复旦大学; 西湖大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视觉-语言-动作模型忽略3D物体结构信息的缺陷,构建概念专家模块生成解析概念,通过双阶段机制提供显式引导,提升了模型的操作成功率与学习效率。

AI 中文摘要

当前视觉-语言-动作(VLA)模型主要依赖2D输入,忽略了3D物理世界中固有的丰富物体结构信息和常识知识,这一缺陷限制了它们的空间感知能力和对复杂、高精度操作的适应性。为弥合这一关键差距,我们为VLA构建了一个概念专家(Concept Expert)模块,以构建可执行的解析概念(Analytic Concepts),将物体表示为显式的、程序化的蓝图。我们的机制在两个协同阶段运行:第一,在VLA推理之前,概念专家利用视觉基础模型(Vision Foundation Models, VFMs)的3D信息来估计初始运动学和结构参数;第二,在整个操作过程中,VLA模型利用其固有能力动态跟踪动态概念参数,不断将其与观测变化对齐,以确保持续的准确性。解析概念一旦建立,就通过以下方式为VLA微调提供显式、高质量的引导:(1)密集的程序化操作奖励;(2)精确的空间引导。这种表述使VLA模型能够学习基于物理的交互行为,同时保持端到端学习的灵活性。我们的实验结果表明,在监督学习和强化学习设置下,成功率和学习效率均实现了持续提升,证明了结构化的、基于概念的引导对VLA后训练的有效性。

英文摘要

Current Vision-Language-Action (VLA) models rely mainly on 2D inputs, neglecting the rich object structural information and commonsense knowledge inherent in the 3D physical world. This deficiency restricts their spatial awareness and adaptability for complex, high-precision manipulation. To bridge this crucial gap, we construct a Concept Expert module for VLA to build executable Analytic Concepts that represent objects as explicit, programmatic blueprints. Our mechanism operates in two synergistic phases: First, prior to VLA inference, the Concept Expert leverages 3D information from Vision Foundation Models (VFMs) to estimate the initial kinematic and structural parameters. Second, throughout the manipulation process, the VLA model utilizes its inherent capability to dynamically track the dynamic concept parameters, continuously aligning them with observational changes to ensure persistent accuracy. Once established, the Analytic Concepts provide explicit, high-quality guidance for VLA fine-tuning through (1) dense, programmatic manipulation rewards and (2) precise spatial guidance. This formulation allows VLA models to learn physically grounded interaction behaviors while maintaining end-to-end learning flexibility. Our experimental results show consistent improvements in success rate and learning efficiency across supervised and reinforcement learning settings, demonstrating the effectiveness of structured, concept-based guidance for VLA post-training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑