Embodied.cpp:面向异构机器人的具身AI模型可移植推理运行时
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots
- Southeast University(东南大学)
- Nanjing University(南京大学)
- Microsoft Research(微软研究院)
- Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院(AIR))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出Embodied.cpp,一个基于C++的可移植推理运行时,通过五层架构支持多速率执行、延迟优先融合推理和可扩展接口,在异构机器人上高效部署VLA和世界动作模型。
AI中文摘要:
具身AI模型现在涵盖视觉-语言-动作(VLA)模型和世界动作模型(WAM),但实际部署仍然分散在特定于模型的Python堆栈、后端假设和机器人端胶水代码中,尤其是在异构边缘设备上。现有的推理运行时主要设计用于请求-响应服务,因此不满足具身部署的运行时契约:闭环控制内的多速率执行、异构硬件上的延迟优先batch-1推理,以及超越固定令牌I/O的可扩展具身接口。我们提出Embodied.cpp,一个用于具身模型的可移植C++推理运行时。基于代表性VLA模型和WAM的架构分析,Embodied.cpp捕获共享执行路径并将其组织为五个层:输入适配器、序列构建器、骨干执行、头部插件和部署适配器。该运行时提供模块化多速率执行、延迟优先融合推理以及可扩展的运算符和I/O支持,通过一个后端抽象实现跨异构设备、机器人和模拟器的部署。我们在两个VLA模型(HY-VLA和pi0.5)以及使用LingBot-VA Transformer块的初步WAM基准上评估Embodied.cpp。VLA部署分别实现了100.0%和91.0%的任务成功率的成功闭环执行。WAM基准将块内存从312.2 MiB减少到88.1 MiB。这些结果表明,Embodied.cpp在保持跨不同具身模型架构的高精度的同时,提高了部署效率。
英文摘要:
Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, and robot-side glue code, especially on heterogeneous edge devices. Existing inference runtimes are designed mainly for request-response serving and therefore do not satisfy the runtime contract of embodied deployment: multi-rate execution inside closed-loop control, latency-first batch-1 inference on heterogeneous hardware, and extensible embodied interfaces beyond fixed token I/O. We present Embodied$.$cpp, a portable C++ inference runtime for embodied models. Based on an architectural analysis of representative VLA models and WAMs, Embodied$.$cpp captures a shared execution path and organizes it into five layers: input adapters, sequence builders, backbone execution, head plugins, and deployment adapters. The runtime provides modular multi-rate execution, latency-first fused inference, and extensible operator and I/O support, enabling deployment across heterogeneous devices, robots, and simulators through one backend abstraction. We evaluate Embodied$.$cpp on three VLA and two WAM models, using normalized comparisons across Python and C++ quantization configurations. Overall, Embodied$.$cpp achieves 1.05x-2.70x inference speedups and 7\%-77\% lower VRAM relative to Python baselines, while maintaining near-baseline success for most configurations. These results show that Embodied$.$cpp improves deployment efficiency while preserving high control quality across diverse embodied model architectures. Project Link: https://github.com/SEU-PAISys/Embodied.cpp