迈向实时VLA:阶段感知的两步流去噪与系统级评估
Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation
浏览论文内容
中文总结 AI 辅助
针对VLA模型推理与机器人执行之间的时间差,提出阶段感知的两步流去噪方法,将推理步数从10降至2,推理时间从61.557ms降至21.956ms,并开发分布式实时框架,在服装折叠任务上验证了有效性。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型面临低速率推理与高速率机器人执行之间的时间差。我们通过模型推理和机器人执行链的端到端延迟测量来表征这一差距。重复的流匹配去噪对推理成本贡献显著,而机器人侧延迟主要源于感知获取、通信调度和物理响应。对速度场的分析显示,在早期积分阶段幅度和方向相对稳定,而在接近终端步骤时出现更强的方向校正。基于这种阶段异质性,我们提出了两阶段非均匀去噪,将步数从10步减少到2步,模型推理时间从61.557毫秒降至21.956毫秒。我们还开发了一个分布式实时VLA框架,具有独立的推理、动作发布和机器人控制速率,模块化观测采集,以及动作来源日志记录。以π0.5为基线,我们在一个长时程物理服装折叠任务上评估了六种实时执行方法。Legato在基于训练的方法中整体表现最佳,而Temporal Smoothing在无训练方法中领先;两者在任务成功率、完成时间、动作连续性和加速度平滑性方面均表现强劲。将两步去噪与代表性执行方法相结合,大幅降低了推理成本,同时任务性能仅有小幅下降。这些结果促使对模型推理效率和机器人系统时序进行联合优化。
英文摘要
Vision-language-action (VLA) models face a timing gap between low-rate inference and high-rate robot execution. We characterize this gap through end-to-end latency measurements of model inference and the robot execution chain. Repeated Flow Matching denoising contributes substantially to inference cost, while robot-side delays mainly arise from perception acquisition, communication scheduling, and physical response. Analysis of the velocity field shows relatively stable magnitude and direction in early integration, followed by stronger directional correction near the terminal steps. Based on this stage heterogeneity, we propose two-stage non-uniform denoising, reducing the number of steps from 10 to 2 and model-inference time from 61.557 ms to 21.956 ms. We also develop a distributed real-time VLA framework with independent inference, action-publication, and robot-control rates, modular observation acquisition, and action-provenance logging. Using π0.5 as the baseline, we evaluate six real-time execution methods on a long-horizon physical garment-folding task. Legato performs best overall among training-based methods, while Temporal Smoothing leads among training-free methods; both perform strongly in task success, completion time, action continuity, and acceleration smoothness. Combining two-step denoising with representative execution methods substantially reduces inference cost with a small reduction in task performance. These results motivate joint optimization of model-inference efficiency and robot-system timing.
发表机构
- Magic-Lab Team, Magiclab Robotics Inc.(Magic-Lab团队,Magiclab机器人公司)
- Zhejiang University(浙江大学)
- Southeast University(东南大学)
- Harbin Institute of Technology(哈尔滨工业大学)
- Tongji University(同济大学)
- Jilin University(吉林大学)
机构由 AI 辅助整理,请以论文原文为准。