发表机构
ZTE Corporation(中兴通讯股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NebulaVLA是一种异步双频VLA模型,通过解耦语义推理与动作控制、引入GESTURE-7表示及引导动作算法,在LIBERO-Plus上以85.5%成功率、约2.7倍速度优于同步基线,实现高效机器人操纵控制。
AI 中文摘要
视觉-语言-动作(VLA)模型的实际部署常受限于效率-性能权衡、跨实体泛化能力及执行流畅度。我们提出NebulaVLA,一种异步双频架构,将高层语义推理与低层动作控制解耦,优化计算资源与模块化。为弥合异构机器人间的语义差距,我们引入GESTURE-7,一种统一的语言接地动作表示;此外,我们的引导动作算法通过基于掩码的流畅度约束强制运动学连续性。综合评估表明,NebulaVLA显著优于同步基线,在LIBERO-Plus上实现85.5%的平均成功率,动作生成速度提升约2.7倍,这种异步设计可为实用机器人提供高效且响应迅速的控制。
英文摘要
Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA, an asynchronous dual-frequency architecture that decouples high-level semantic reasoning from low-level action control, optimizing computational resources and modularity. To bridge semantic gaps across heterogeneous robots, we introduce GESTURE-7, a unified language-grounded action representation. Furthermore, our Guide Action algorithm enforces kinematic continuity via mask-based smoothness constraints. Comprehensive evaluations demonstrate that NebulaVLA significantly outperforms synchronous baselines, achieving an 85.5\% average success rate on LIBERO-Plus and accelerating action generation by \textasciitilde 2.7$\times$. This asynchronous design enables highly efficient and responsive control for practical robotics.
Comments14 pages, 5 figures