arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13926cs.RO

S平方-VLA:自动驾驶视觉-语言-动作模型中语义与空间流的解耦

S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving

Jianguo Yu, Rukang Wang, Duanfeng Chu, Chen Wang, Renju Feng, Liping Lu

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对自动驾驶中视觉语言模型生成低级控制动作的局限,提出S平方-VLA解耦语义和空间流,语义流用于意图推理,空间流保留空间特征并赋予先验,双流规划适配器融合二者,在基准测试中取得新的最先进水平,优于基线。

中文摘要 AI 辅助

视觉语言模型(VLMs)在自动驾驶高级推理中潜力显著,但在生成精确低级控制动作上存在局限,源于离散语言令牌与连续轨迹规划的不匹配导致的语义-物理差距。视觉语言动作(VLA)架构试图弥合差距,却造成新瓶颈,标准VLA存在空间表示崩溃。为此提出S平方-VLA,解耦语义和空间流。语义流利用分层桥接提取多尺度VLM特征进行意图推理,空间流绕过自回归语言瓶颈,保留视觉编码器的未压缩空间特征,通过辅助感知监督赋予模型丰富空间和几何先验,双流规划适配器融合语义意图与空间约束。在NAVSIM闭环基准测试中,S平方-VLA在纯监督微调设置下取得新的VLA模型最先进水平,缓解了传统VLMs的空间表示崩溃,显著优于基线。

英文摘要

Vision-Language Models (VLMs) have demonstrated remarkable potential for high-level reasoning in autonomous driving, yet they fundamentally struggle to generate precise, low-level control actions. This limitation is rooted in a semantic-physical gap caused by the inherent mismatch between discrete language tokens and continuous trajectory planning. While Vision-Language-Action (VLA) architectures attempt to bridge this gap by unifying perception and control into a single policy, this entanglement creates a new bottleneck. Standard VLAs experience a severe spatial representation collapse, which irreversibly degrades the fine-grained spatial and geometric priors essential for safe, boundary-aware navigation. To address this limitation, we propose the S-squared-VLA, which explicitly decouples the semantic and spatial streams in Vision-Language-Action models. The semantic stream leverages hierarchical bridging to extract multi-scale VLM features for robust intent reasoning. In parallel, an independent spatial stream bypasses the autoregressive language bottleneck, directly preserving uncompressed spatial features from the visual encoder. By integrating auxiliary perception supervision, this stream explicitly equips the model with rich spatial and geometric priors. Finally, a dual-stream planning adapter fuses high-level semantic intent with precise spatial constraints via cascaded attention mechanisms. Evaluations on the NAVSIM closed-loop benchmark show that S-squared-VLA achieves a Predictive Driver Model Score (PDMS) of 87.1, establishing a new state-of-the-art for VLA models under a purely supervised fine-tuning (SFT) setting. By mitigating the spatial representation collapse of traditional VLMs, our framework significantly outperforms baselines, achieving the highest No Collision (NC) rate of 98.4 among all evaluated methods.

发表机构

  • School of Mechanical and Electronic Engineering, Wuhan University of Technology(武汉理工大学机电工程学院)
  • Intelligent Transportation Systems Research Center, Wuhan University of Technology(武汉理工大学智能交通系统研究中心)
  • School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

↑