AI 中文总结
针对移动机器人VLA部署的时延冲突,提出CloudEdgeVLA云边协同策略,通过表征学习处理时间错位,在LIBERO套件上远优于基线,实现高时延下的高任务成功率。
AI 中文摘要
将拥有数十亿参数的视觉-语言-动作(VLA)策略部署在移动机器人上会产生系统层面的冲突:语义推理依赖云端GPU,而闭环控制必须在网络延迟和抖动的情况下实现本地响应。现有的分层异步策略虽能提升吞吐量,但其慢路径表征仍可能存在过时问题,或需要明确的调度和延迟提示。本文提出CloudEdgeVLA,一种将时间错位视为表征学习问题的云边策略:云端VLA将延迟观测编码为缓慢变化的任务特征,轻量的边缘头则结合最新可用的云端特征与当前本地视觉信息。训练时,当前帧和随机延迟帧分别与新鲜路径和过时路径中的同一当前动作目标配对,该目标促使云端表征保留任务级信息,同时边缘路径提供状态敏感的修正。在四个LIBERO套件上,CloudEdgeVLA在40步均匀延迟窗口下仍保持63.8%至78.0%的成功率,而VLASH最高仅达6.4%,评估的单路径基线最高仅达3.0%。通过移除控制环路中的阻塞同步,该设计为可扩展VLA部署提供了实用途径,使云端模型可扩展,同时边缘计算保持轻量且响应迅速。
英文摘要
Deploying billion-parameter Vision-Language-Action (VLA) policies on mobile robots creates a systems conflict: semantic reasoning benefits from cloud GPUs, whereas closed-loop control must respond locally despite network delay and jitter. Existing hierarchical and asynchronous policies improve throughput, but their slow-path representations can still arrive stale or require explicit scheduling and delay cues. We introduce CloudEdgeVLA, a cloud-edge policy that treats temporal misalignment as a representation-learning problem. A cloud VLA encodes delayed observations into slowly varying task features, while a lightweight edge head combines the latest available cloud feature with current local vision. During training, current and randomly delayed frames are paired with the same current action target in fresh and stale paths. This objective encourages the cloud representation to preserve task-level information while the edge path supplies state-sensitive corrections, driving emergent specialization. Across four LIBERO suites, CloudEdgeVLA retains 63.8-78.0% success with a 40-step uniform-delay window, whereas VLASH reaches at most 6.4% and the evaluated single-path baselines at most 3.0%. By removing blocking synchronization from the control loop, the design offers a practical route to scalable VLA deployment in which cloud models can grow while edge computation remains lightweight and responsive.