arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于Mamba的选择性状态空间建模改善SmolVLA视觉-语言-动作专家模型的精度-复杂度权衡

Mamba-based Selective State Space Modeling Improves the Accuracy-Complexity Tradeoff of SmolVLA Vision-Language-Action Experts

Farida Mohsen, Thowayba Elkaffash, Mohammad Reza Chalak Qazani, Mohamed Mabrok, Nader Meskin, Ali Safa

arXiv 2608.21407首次发表:更新:

AI 中文总结

本文将SmolVLA的动作专家模块替换为Mamba的选择性状态空间建模,在LIBERO基准上,于不同执行范围下均优于Transformer基线,在N=1时参数复杂度降低24%,N=50时性能提升7.8%。

AI 中文摘要

视觉-语言-动作(VLA)模型面临任务成功率与策略调用频率之间的关键权衡:每次推理执行单个动作(N=1)可实现准确的机器人控制,但会产生巨大的计算时间开销,导致无法实现实时部署;而在重新规划前执行更长的动作范围(N≫1)则会降低计算复杂度,但不可避免地会降低系统的成功率。为改善VLA模型的精度-复杂度权衡,本文将流行的SmolVLA模型(以高精度且低复杂度作为参考模型)的动作专家模块中的因果自注意力替换为Mamba的选择性状态空间建模。我们在广泛使用的LIBERO基准套件上,针对三个执行范围N∈{1,25,50}(分别对应高、中、低计算复杂度),分别评估基于Mamba和Transformer的专家模型。结果显示,Mamba专家的优势随执行范围增大而提升,在长执行范围N=50和N=25下均能显著保持成功率:当在重新规划前执行N=50个动作(对应可行的实时部署)时,Mamba专家比Transformer基准模型高出7.8%;当N=25时,Mamba专家比Transformer基准模型高出3.7%;在按动作重新规划(N=1)时,Mamba变体与基于Transformer的平均成功率相当,且凭借Mamba的高效计算特性,整体模型参数复杂度显著降低24%。

英文摘要

Vision-language-action (VLA) models face a crucial tradeoff between their task success rate and the policy-call frequency. Executing a single action per inference ($N=1$) enables accurate robot control but comes at the cost of huge compute time overheads, making real-time implementation infeasible. On the other hand, executing longer action horizons before replanning ($N\gg1$) reduces compute complexity, but inevitably degrades the system's success rate. In order to improve the VLA accuracy-complexity tradeoff, this paper investigates Mamba's selective state-space modeling as an alternative to causal self-attention within the action expert of the popular SmolVLA model, widely used as a reference model for its highly accurate yet low complexity nature. We evaluate both the Mamba- and Transformer-based experts on the widely-adopted LIBERO benchmark suites across three execution horizons $N\!\in\!\{1,25,50\}$, respectively corresponding to high, moderate and low compute complexities. Our results remarkably show that the advantage of the Mamba expert increases with the execution horizon, indicating significant success retention under long execution horizons $N = 50$ and $N = 25$. When $N = 50$ actions are executed before replanning (i.e., corresponding to feasible real-time deployment), the Mamba expert outperforms the Transformer baseline by $7.8\%$. In addition, when $N = 25$ actions are executed before replanning, our Mamba expert outperforms the Transformer baseline by $3.7\%$. Finally, under per-action replanning ($N=1$), our Mamba variant matches the Transformer-based mean success rate while significantly reducing the overall model parameter complexity by $24\%$ thanks to Mamba's compute-efficient nature.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑