ComVLA:6G互联机器人中VLA模型的通信感知分割推理
ComVLA: Communication-Aware Split Inference for VLA Models in 6G-Connected Robotics
- Technical University of Berlin(柏林工业大学)
- Huawei Heisenberg Research Center(华为海森堡研究中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对6G互联机器人中VLA模型云推理受无线信道限制的问题,提出利用语言指导自适应调整令牌预算的ComVLA框架,在LIBERO上以少量成功率代价大幅降低计算和延迟。
AI中文摘要:
互联机器人是一种新兴的6G应用,其中移动机器人遵循自然语言指令来操作物理对象。实现这一功能的视觉-语言-动作(VLA)模型规模过大,无法在机器人上运行;常见的趋势是将推理卸载到云端。然而,无线链路限制了边缘在每个控制步骤中可传输的传感数据量。最近的两条研究路线解决了这一约束:语义通信编解码器压缩传感器数据,但需要针对特定信道进行重新训练;VLA令牌剪枝器从图像中选择令牌,但忽略了信道。我们的洞察是,语言中包含的密集语义信息已经指示了哪些视觉令牌是重要的。我们提出了ComVLA,一个利用这种语言指导来使VLA令牌预算适应信道容量的框架。在LIBERO基准上,传输32个令牌而不是512个,与原始OpenVLA-OFT基线相比,ComVLA将推理计算减少了74%,推理延迟减少了22%,代价是平均任务成功率下降1.5个百分点(95.4%对96.9%),并且在瑞利和莱斯衰落条件下保持在容量预算内。这些结果表明,协同设计VLA推理和无线通信是6G互联机器人的一个实用方向。
英文摘要:
Connected robotics is an emerging 6G application where mobile robots follow natural-language instructions to manipulate physical objects. The Vision-Language-Action (VLA) models that enable this are too large to run on the robot; a common trend is to offload inference to the cloud. The wireless link, however, limits how much sensing data the edge can transmit per control step. Two recent lines address this constraint: semantic communication codecs compress sensor data but require channel-specific retraining, and VLA token pruners select tokens from image but ignore the channel. Our insight is that the dense semantic information contained in the language already indicates which visual tokens matter. We propose ComVLA, a framework that uses this language guidance to adapt the VLA token budget to the channel capacity. Transmitting 32 tokens instead of 512 on the LIBERO benchmark, ComVLA cuts inference compute by 74% and inference latency by 22% versus the original OpenVLA-OFT baseline, at a cost of 1.5 pp in average task success (95.4% vs. 96.9%), and it stays within the capacity budget under Rayleigh and Rician fading. These results demonstrate that co-designing VLA inference and wireless communication is a practical direction for 6G-connected robotics.