sensVLA:面向自主轮式装载机的空间接地视觉-语言-动作模型
sensVLA: Spatially-Grounded Vision-Language-Action Model for Autonomous Wheel Loader
浏览论文内容
中文总结 AI 辅助
本文提出sensVLA,一种结合Qwen3-2B VLM与流匹配动作专家的VLA架构,通过BEV特征交叉注意力实现空间接地,在轮式装载机数据集上降低纵向速度RMSE 28%并增强容错性。
中文摘要 AI 辅助
自主轮式装载机控制需要对任务语义、自我中心视觉、本体感觉和3D场景几何进行联合推理。我们提出sensVLA,一种视觉-语言-动作(VLA)架构,它结合了Qwen3-2B视觉-语言模型(VLM)与一个完全可训练的、通过流匹配速度回归训练的Transformer动作专家。sensVLA通过专用的交叉注意力通路,将从融合的前后激光雷达提取的鸟瞰图(BEV)特征直接路由到动作专家,而VLM则消费前后RGB视图以提供任务条件化的语义上下文。这种设计将空间接地与语言推理解耦,同时在决策时保留两个流之间的交互。该专家预测六个动作维度:纵向速度、转向、车身框架位移、臂速和铲斗速率。在来自轮式装载机的真实世界数据集上,sensVLA达到了与强相机基线相当的整体每步性能,并在装载中心场景中将纵向速度RMSE降低了28%,位移误差降低了9%。当相机流被损坏或移除时,它退化程度也减少了29%,表明显式的空间接地提高了重型设备自主性的准确性和容错性。
英文摘要
Autonomous wheel-loader control requires joint reasoning over task semantics, egocentric vision, proprioception, and 3D scene geometry. We present sensVLA, a Vision-Language-Action (VLA) architecture that combines a Qwen3-2B Vision-Language Model (VLM) with a fully trainable transformer action expert trained by flow-matching velocity regression. sensVLA routes Bird's-Eye-View (BEV) features, extracted from fused front and rear lidar, directly to the action expert through a dedicated cross-attention pathway, while the VLM consumes front and rear RGB views to provide task-conditioned semantic context. This design decouples spatial grounding from linguistic reasoning while preserving interaction between both streams at decision time. The expert predicts six action dimensions: longitudinal velocity, steering, body-frame displacement, arm rate, and bucket rate. On a real-world dataset from a wheel loader, sensVLA reaches aggregate per-step parity with a strong camera-only baseline and reduces longitudinal velocity RMSE by 28% and displacement error by 9% on loading centric scenarios. It also degrades 29% less when the camera stream is corrupted or removed, evidencing that explicit spatial grounding improves accuracy and fault-tolerance for heavy equipment autonomy.
发表机构
- sensmore GmbH(sensmore 有限公司)
机构由 AI 辅助整理,请以论文原文为准。