算子不匹配问题:基于便携式GPU计算部署BEV感知
The Operator Mismatch Problem: Deploying BEV Perception with Portable GPU Compute
浏览论文内容
中文总结 AI 辅助
针对BEV感知模型部署的算子不匹配问题,提出便携式GPU计算框架BEVPIPE,实现19.5倍端到端加速且保留98.5% mAP,支持跨GPU后端移植。
中文摘要 AI 辅助
现代自动驾驶系统依赖鸟瞰视角(BEV)感知模型,这类模型融合相机与激光雷达(LiDAR)输入以检测3D空间中的物体,精度较高,但无法通过标准推理运行时部署。原因在于存在算子不匹配问题:密集卷积(运行时可良好处理)、稀疏3D卷积(运行时无法表示)以及几何散射操作(运行时无对应实现)之间不兼容。当前所有稀疏卷积库均仅支持CUDA且与PyTorch绑定,导致BEV部署被锁定在单一厂商硬件与单一执行框架中。本文提出BEVPIPE框架,该框架利用便携式GPU计算API部署多模态BEV感知流水线,并将其与生产推理运行时集成。BEVPIPE将模型划分为运行时管理的密集子图,以及三个外部算子扩展模块(体素化器、稀疏编码器、BEV投影器),各模块通过共享GPU内存空间连接。BEVPIPE相比传统部署实现了19.5倍的端到端加速,同时保留了98.5%的参考平均精度均值(mAP),且可跨不同GPU后端移植。
英文摘要
Modern autonomous driving systems rely on bird's-eye-view (BEV) perception models that fuse camera and LiDAR inputs to detect objects in 3D space. These models are accurate, but they cannot be deployed through standard inference runtimes. The reason is an operator mismatch between dense convolutions (which runtimes handle well), sparse 3D convolutions (which runtimes cannot represent), and geometric scatter operations (which runtimes have no vocabulary for). Today, every sparse convolution library is CUDA-only and PyTorch-coupled, locking BEV deployment to a single vendor's hardware and a single execution framework. We present BEVPIPE, a framework for deploying multimodal BEV perception pipelines using portable GPU compute APIs and integrating them with production inference runtimes. BEVPIPE partitions the model into runtime-managed dense subgraphs and three external operator extensions (voxelizer, sparse encoder, BEV projector), connected through a shared GPU memory space. BEVPIPE achieves a 19.5x end-to-end speedup over conventional deployments while retaining 98.5% of reference mAP. We also showcase that BEVPIPE is portable across different GPU backends.
发表机构
- Intel Corporation(英特尔公司)
机构由 AI 辅助整理,请以论文原文为准。