发表机构
Yeungnam University; Kyungpook National University(岭南大学; 庆北国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于Mamba的VLA模型GaussVLA,通过高斯空间分词器和深度感知思维链提升空间推理能力,在LIBERO数据集上仅2亿参数就实现93.5%平均成功率,优于SpatialVLA且参数效率更高。
AI 中文摘要
视觉-语言-动作(VLA)模型将视觉观测编码为扁平的2D patch token,这类token不携带内在几何结构;为其添加密集单目深度信息时,仅注入了每个像素的标量值,既未编码表面朝向也未编码几何置信度,这使得策略在动作预测时的结构化空间推理能力受限。我们提出GaussVLA,这是一种基于Mamba的VLA模型,包含两个定制模块:高斯空间分词器(GST),用于将冻结的语义和深度特征提升为紧凑的3D高斯token,并用学习到的查询池化几何显著区域;以及深度感知思维链(DA-CoT),该模块在语言和流时间条件下执行结构化、非自回归的几何推理。在模拟和真实世界评估中,GaussVLA展现出强大的空间操作性能,同时保持参数高效性:在LIBERO数据集上,它仅用2亿参数就达到了93.5%的平均成功率,在Spatial套件上实现了100.0%的成功率,相对平均成功率比SpatialVLA提升了19.7%,且参数效率显著更高。
英文摘要
Vision-Language-Action (VLA) models encode visual observations as flat 2D patch tokens that carry no intrinsic geometric structure, and augmenting them with dense monocular depth injects per-pixel scalar values that encode neither surface orientation nor geometric confidence. This leaves the policy with limited structured spatial reasoning for action prediction. We propose GaussVLA, a Mamba-based VLA that incorporates two custom modules: Gaussian Spatial Tokenizer (GST) to lift frozen semantic and depth features into compact 3D Gaussian tokens, pools geometrically salient regions with learned queries, and \emph{Depth-Aware Chain-of-Thought (DA-CoT)} that performs structured, non-autoregressive geometric reasoning under language and flow-time conditioning. Across both simulation and real-world evaluations, GaussVLA demonstrates strong spatial-manipulation performance while remaining parameter-efficient. On LIBERO, it achieves 93.5% average success and 100.0% success on the Spatial suite with only 200M parameters, improving over SpatialVLA by 19.7% relative average success while remaining significantly more parameter-efficient.
CommentsAccepted to BMVC 2026