arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03682cs.AIcs.RO

PhyAI:边缘端实时物理人工智能,云端可扩展部署

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud

Chenghua Wang, Daliang Xu, Dongqi Cai, Duojin Sun, Hao Zhang, Haoze Qian, Huaiyuan Zhang, Jinshuo Cui, Junbo Cui, Kezhao Zhao, Longxi Gao, Mengwei Xu, Rongjie Y… 展开作者

Chenghua Wang, Daliang Xu, Dongqi Cai, Duojin Sun, Hao Zhang, Haoze Qian, Huaiyuan Zhang, Jinshuo Cui, Junbo Cui, Kezhao Zhao, Longxi Gao, Mengwei Xu, Rongjie Yi, Ruixin Liu, Shangguang Wang, Tam Sikyuen, Tianyue Zhang, Weikai Xie, Xuanzhe Liu, Yingying Qin, Yiwen Lu, Yuan Yao, Yuezhi Zu, Yunhan Guo, Yuxin Zheng, Ziqi Guo

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对物理AI多部署场景推理程序不统一的问题,提出PhyAI推理引擎,实现多模型跨端部署,在多个基准上取得1.40-4.65倍加速,还引入控制时间屋顶线分析模型特性。

中文摘要 AI 辅助

物理人工智能(Physical AI)策略在其整个生命周期中都需要进行推理,包括模型评估、云端强化学习部署、边缘GPU服务以及机载部署。尽管这些设置共享相同的检查点和动作语义,但它们通常依赖于独立的推理程序。为了统一这些设置,我们构建了PhyAI,这是一个物理人工智能推理引擎,拥有单一运行时环境,该环境将特定于架构的条件设置、求解器、缓存和输出逻辑保留在模型适配器中,同时共享图执行、内核、内存管理和并行服务。相同的代码库可在机载、边缘和云端部署的单个或多个GPU上运行视觉-语言-动作(VLA)模型和世界-动作模型(WAMs)。我们在MiniCPM-Robot发布当天就通过适配器接口将其添加到PhyAI中。PhyAI在pi0、pi0.5、GR00T N1.7和MiniCPM-Robot的官方实现上实现了1.40倍至4.65倍的加速。在Cosmos3-Nano-Policy-DROID上,它在8个H20 GPU(CFG=2,TP=4)上将延迟从2.46秒降低到1.18秒,实现了2.08倍的加速。在多种配置下,专用运行时仍然更快,因此我们的目标是拥有一个具有竞争力延迟的单一运行时,而非在每种情况下都实现最快结果。详细分析显示了不同模型需要不同执行策略的原因:在批大小为1的Hopper系列GPU上,pi0.5动作专家占8.8%的浮点运算量(FLOPs),但占57.2%的延迟;在批大小为32时,其占比降至13.5%,吞吐量达到约100个样本/秒。Cosmos3仍然以生成为主导,当批大小从1增加到16时,其吞吐量仅增加14.3%。我们进一步引入了控制时间屋顶线(control-time Roofline),该方法区分推理受限和环境受限的控制;在四个LIBERO套件上测量的pi0.5点是环境受限的,而Cosmos3保持推理受限。代码和基准:此https URL。

英文摘要

Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, they often rely on separate inference programs. To unify them, we build PhyAI, a Physical AI inference engine with a single runtime that keeps architecture-specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services. The same codebase runs vision-language-action (VLA) models and world-action models (WAMs) on single or multiple GPUs across onboard, edge, and cloud deployments. We used the adapter interface to add MiniCPM-Robot on the day of its release. PhyAI achieves 1.40x-4.65x speedups over the official implementations of pi0, pi0.5, GR00T N1.7, and MiniCPM-Robot. On Cosmos3-Nano-Policy-DROID it reduces latency from 2.46 to 1.18 s on eight H20 GPUs (CFG=2, TP=4), a 2.08x speedup. Specialized runtimes remain faster in several configurations, so our goal is one runtime with competitive latency rather than the fastest result in every case. Detailed profiles reveal why different models need different execution policies: on a Hopper-series GPU at batch size one, the pi0.5 action expert accounts for 8.8% of FLOPs but 57.2% of latency; at batch size 32 its share drops to 13.5% and throughput reaches about 100 samples/s. Cosmos3 remains generation-dominated and gains only 14.3% throughput as batch size increases from 1 to 16. We further introduce the control-time Roofline, which distinguishes inference-bound from environment-bound control; the measured pi0.5 points on four LIBERO suites are environment-bound while Cosmos3 stays inference-bound. Code and benchmarks: https://github.com/mingti-org/phyai.

补充信息

↑