arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.02501cs.ROcs.CVcs.OS

Embodied.cpp:面向异构机器人的具身AI模型可移植推理运行时

Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

  • Southeast University(东南大学)
  • Nanjing University(南京大学)
  • Microsoft Research(微软研究院)
  • Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院(AIR))

机构由 AI 辅助整理,请以论文原文为准。

Ling Xu, Borui Li, Hao Wu, Chuyu Han, Xiangyu Li, Mohan Hua, Shiqi Jiang, Ting Cao, Chuanyou Li, Sheng Zhong, Shuai Wang

AI总结:

提出Embodied.cpp,一个基于C++的可移植推理运行时,通过五层架构支持多速率执行、延迟优先融合推理和可扩展接口,在异构机器人上高效部署VLA和世界动作模型。

AI中文摘要:

具身AI模型现在涵盖视觉-语言-动作(VLA)模型和世界动作模型(WAM),但实际部署仍然分散在特定于模型的Python堆栈、后端假设和机器人端胶水代码中,尤其是在异构边缘设备上。现有的推理运行时主要设计用于请求-响应服务,因此不满足具身部署的运行时契约:闭环控制内的多速率执行、异构硬件上的延迟优先batch-1推理,以及超越固定令牌I/O的可扩展具身接口。我们提出Embodied.cpp,一个用于具身模型的可移植C++推理运行时。基于代表性VLA模型和WAM的架构分析,Embodied.cpp捕获共享执行路径并将其组织为五个层:输入适配器、序列构建器、骨干执行、头部插件和部署适配器。该运行时提供模块化多速率执行、延迟优先融合推理以及可扩展的运算符和I/O支持,通过一个后端抽象实现跨异构设备、机器人和模拟器的部署。我们在两个VLA模型(HY-VLA和pi0.5)以及使用LingBot-VA Transformer块的初步WAM基准上评估Embodied.cpp。VLA部署分别实现了100.0%和91.0%的任务成功率的成功闭环执行。WAM基准将块内存从312.2 MiB减少到88.1 MiB。这些结果表明,Embodied.cpp在保持跨不同具身模型架构的高精度的同时,提高了部署效率。

英文摘要:

Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, and robot-side glue code, especially on heterogeneous edge devices. Existing inference runtimes are designed mainly for request-response serving and therefore do not satisfy the runtime contract of embodied deployment: multi-rate execution inside closed-loop control, latency-first batch-1 inference on heterogeneous hardware, and extensible embodied interfaces beyond fixed token I/O. We present Embodied$.$cpp, a portable C++ inference runtime for embodied models. Based on an architectural analysis of representative VLA models and WAMs, Embodied$.$cpp captures a shared execution path and organizes it into five layers: input adapters, sequence builders, backbone execution, head plugins, and deployment adapters. The runtime provides modular multi-rate execution, latency-first fused inference, and extensible operator and I/O support, enabling deployment across heterogeneous devices, robots, and simulators through one backend abstraction. We evaluate Embodied$.$cpp on three VLA and two WAM models, using normalized comparisons across Python and C++ quantization configurations. Overall, Embodied$.$cpp achieves 1.05x-2.70x inference speedups and 7\%-77\% lower VRAM relative to Python baselines, while maintaining near-baseline success for most configurations. These results show that Embodied$.$cpp improves deployment efficiency while preserving high control quality across diverse embodied model architectures. Project Link: https://github.com/SEU-PAISys/Embodied.cpp

补充信息

↑