arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NebulaVLA:用于机器人操纵的带引导动作的双频视觉-语言-动作模型

NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation

Cong Zhao, Shuai Tian, Xu Zhang, Baocheng Ni, Xinguo Song, Xueying Sun, Shu Jiang, Shouchang Yang, Bo Tang, Jin Deng, Ge Zhu, YongCheng Wang, Jin Xu, Ri Yang

arXiv 2608.16503首次发表:更新:

发表机构

ZTE Corporation(中兴通讯股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

NebulaVLA是一种异步双频VLA模型,通过解耦语义推理与动作控制、引入GESTURE-7表示及引导动作算法,在LIBERO-Plus上以85.5%成功率、约2.7倍速度优于同步基线,实现高效机器人操纵控制。

AI 中文摘要

视觉-语言-动作(VLA)模型的实际部署常受限于效率-性能权衡、跨实体泛化能力及执行流畅度。我们提出NebulaVLA,一种异步双频架构,将高层语义推理与低层动作控制解耦,优化计算资源与模块化。为弥合异构机器人间的语义差距,我们引入GESTURE-7,一种统一的语言接地动作表示;此外,我们的引导动作算法通过基于掩码的流畅度约束强制运动学连续性。综合评估表明,NebulaVLA显著优于同步基线,在LIBERO-Plus上实现85.5%的平均成功率,动作生成速度提升约2.7倍,这种异步设计可为实用机器人提供高效且响应迅速的控制。

英文摘要

Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA, an asynchronous dual-frequency architecture that decouples high-level semantic reasoning from low-level action control, optimizing computational resources and modularity. To bridge semantic gaps across heterogeneous robots, we introduce GESTURE-7, a unified language-grounded action representation. Furthermore, our Guide Action algorithm enforces kinematic continuity via mask-based smoothness constraints. Comprehensive evaluations demonstrate that NebulaVLA significantly outperforms synchronous baselines, achieving an 85.5\% average success rate on LIBERO-Plus and accelerating action generation by \textasciitilde 2.7$\times$. This asynchronous design enables highly efficient and responsive control for practical robotics.

Comments14 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑