arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StrataVLA:面向视觉-语言-动作模型的分层高效3D几何基础

StrataVLA: Hierarchical and Efficient 3D Geometric Grounding for Vision-Language-Action Models

Jin Cui, Zhaoyu Pu, Botao Cai, Jun Ye, Xinyue Long, Boran Zhao, Pengju Ren

arXiv 2609.26071首次发表:更新:

AI 中文总结

StrataVLA提出分层几何基础框架,通过冻结几何模型和稀疏适配器向VLA骨干注入3D空间信息,结合任务路由与LRU缓存,在LIBERO达98.53%成功率并减少88%几何调用。

AI 中文摘要

视觉-语言-动作(VLA)模型从大规模视觉-语言预训练中继承了强大的语义先验,但在机器人操作中仍因3D空间感知不足而受限。现有方法要么需要显式的深度或点云输入,要么将几何信息压缩为训练时的监督信号,要么仅在模型输入或动作专家处注入几何信息,使得视觉-语言骨干网络无法持续访问与任务相关的空间信息。我们提出StrataVLA,一种用于分层几何基础(grounding)的即插即用框架。一个冻结的几何基础模型从RGB观测中提取共享的几何特征,而稀疏的、逐层的几何适配器(Geometry Adapters)允许所选骨干网络深度处的视觉表征通过交叉注意力检索相关的几何证据。为了使推理时的几何处理切实可行,StrataVLA进一步将任务感知路由与利用任务操作期间时间冗余的LRU特征缓存相结合。在LIBERO、SimplerEnv和真实世界操作上的实验表明,与强VLA基线相比,StrataVLA取得了持续的性能提升。StrataVLA在LIBERO套件上实现了98.53%的平均成功率,同时将几何模型调用次数减少了高达88%,确立了分层几何注入作为实现空间基础机器人控制的有效且高效的方法。

英文摘要

Vision-Language-Action (VLA) models inherit strong semantic priors from large-scale vision-language pretraining, yet remain limited in robotic manipulation by insufficient 3D spatial awareness. Existing approaches either require explicit depth or point-cloud inputs, compress geometry into training-time supervision, or inject it only at the model input or action expert, leaving the vision-language backbone without persistent access to task-relevant spatial information. We introduce StrataVLA, a plug-and-play framework for hierarchical geometric grounding. A frozen geometry foundation model extracts shared geometric features from RGB observations, while sparse, layer-specific Geometry Adapters allow visual representations at selected backbone depths to retrieve relevant geometric evidence through cross-attention. To make inference-time geometry practical, StrataVLA further combines task-aware routing with an LRU feature cache that exploits temporal redundancy during task manipulation. Experiments on LIBERO, SimplerEnv, and real-world manipulation demonstrate consistent gains over strong VLA baselines. StrataVLA achieves 98.53% average success on LIBERO suites while reducing geometry-model invocations by up to 88%, establishing hierarchical geometry injection as an effective and efficient way to achieve spatially grounded robotic control.

Comments11 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑