arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具身智能业务算子NPU片内融合部署与优化研究

Research on Intra-Chip Fusion Deployment and Optimization of Embodied Intelligence Business Operator NPU

Yuchen Zhu, Longxiang Yin, Wanyu Wang, Jieke Lin, Guoqiang Zou, Zirui Cao, Yuling Yuan, Xiaolan Fan, Lifen Chen, Hao Zheng, Qizhang He, Hongyu Zhou, Chunhai Yu

arXiv 2609.05424首次发表:更新:

发表机构

Beijing Information Science and Technology University; Institute of Computing Technology, Chinese Academy of Sciences; ShanghaiTech University; Hangzhou Institute for Advanced Study, UCAS(北京信息科技大学; 中国科学院计算技术研究所; 上海科技大学; 中国科学院大学杭州高等研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对具身智能感知-计算-控制任务分离部署的延迟与效率问题,本文在飞腾-寒武纪国产异构平台上提出NPU片内融合部署与全流程算子协同优化策略,实现MLU370单卡集成执行,延迟18.7ms,较Jetson AGX Xavier加速2.89倍。

AI 中文摘要

具身智能计算融合了感知、计算与控制。传统上将这三类任务分开部署会导致频繁的数据传输、高延迟和低硬件效率,无法满足动态场景下毫秒级的实时性要求。此外,大多数算子优化方法依赖国外GPU平台,而对国产飞腾-寒武纪异构架构的全流程协同优化尚不充分。本文构建了以飞腾FT-2000/4处理器和寒武纪MLU370加速卡为核心的国产异构计算平台,并针对感知、计算与控制流水线提出了一种NPU片内融合部署与全流程算子协同优化策略。面向具身机器人应用,开展了模块化优化,包括用于高速成像的运动模糊校正算子的MLU硬件适配、ViT模型的轻量级推理优化,以及面向多自由度逆运动学求解的定制算子开发。提出了一种基于片内数据闭环和流水线协作的单卡方案,在MLU370上实现了所有感知-计算-控制任务的集成执行。实验结果表明,所提方法实现了18.7毫秒的全流程单帧延迟,相比NVIDIA Jetson AGX Xavier获得了2.89倍的加速,MLU利用率达82.6%,且精度与主流平台相当。本工作为集成式具身智能计算服务的国产化工程应用提供了实用参考。

英文摘要

Embodied intelligent computing integrates perception, computation and control. Traditional separate deployment of the three tasks leads to frequent data transmission, high latency and low hardware efficiency, failing to satisfy millisecond-level real-time requirements in dynamic scenarios. Besides, most operator optimization methods rely on foreign GPU platforms, while full-process collaborative optimization for domestic Phytium-Cambricon heterogeneous architectures is insufficient. This paper builds a domestic heterogeneous computing platform with Phytium FT-2000/4 processor and Cambricon MLU370 acceleration card, and proposes an NPU on-chip fusion deployment and full-process operator collaborative optimization strategy for perception, computation and control pipelines. Targeting embodied robot applications, modular optimization is conducted, including MLU hardware adaptation of motion blur correction operators for high-speed imaging, lightweight inference optimization of ViT models, and customized operator development for multi-DOF inverse kinematics solution. An on-chip data closed-loop and pipeline collaboration-based single-card solution is proposed to implement integrated execution of all perception-computation-control tasks on MLU370. Experimental results show that the proposed method achieves a full-process single-frame latency of 18.7 ms and a speedup of 2.89 compared with NVIDIA Jetson AGX Xavier, with 82.6% MLU utilization and comparable accuracy to mainstream platforms. This work offers a practical reference for domestic engineering applications of integrated embodied intelligent computing services.

Comments25 pages, 16 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑