arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FlashBEV:具有I/O感知的快速且内存高效的精确BEV变换

FlashBEV: Fast and Memory-Efficient Exact BEV Transformation with IO-Awareness

Shunsuke Yokokawa, Hironori Kasahara

arXiv 2607.10071首次发表:更新:

发表机构

Waseda University; T2, Inc.(早稻田大学; T2公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究自动驾驶中BEV感知的视图变换问题,提出FlashBEV执行策略,通过重新设计执行方式,使其在数学上等同于Tensorized Sampling-VT,减少内存流量和开销,降低峰值GPU内存,加速推理延迟,突破可扩展性障碍。

AI 中文摘要

鸟瞰(BEV)感知是自动驾驶中基于相机的3D理解的核心组成部分,其中视图变换(VT)将多相机图像特征映射到统一的BEV表示中。基于采样的视图变换(Sampling-VT)很有吸引力,因为它支持用于高分辨率和远距离感知的密集且连续的BEV聚合。然而,其部署瓶颈在于系统层面:标准的张量化Sampling-VT实现(称为Tensorized Sampling-VT)会显式地生成与高度相关的大型中间张量,导致内存和延迟成本随垂直分辨率和相机数量扩展不佳。我们从算子执行的角度重新审视Tensorized Sampling-VT,发现它遵循收集-归约模式:每个BEV查询独立地在相机和高度区间上累积贡献,实现线程本地累积和即时重新计算,无需生成与高度和相机相关的中间结果。基于此,我们提出了FlashBEV,这是一种完全融合且具有I/O感知的执行策略,在数学上等同于Tensorized Sampling-VT(相同的算子输出),同时大幅减少全局内存流量和内核启动开销。实验表明,FlashBEV的峰值GPU内存降低了一个数量级以上,推理延迟显著加速,内存实际上与高度区间数量无关,将算子的峰值内存降低到O(BCXY)(仅输出)。这在内存受限设备上的固定部署预算内解锁了更高的BEV范围/分辨率和垂直离散化。我们的贡献是一种执行重新设计——相同的数学,不同的执行方式——消除了可部署的Sampling-VT的关键可扩展性障碍。代码可在这个https URL获取。

英文摘要

Bird's-eye-view (BEV) perception is a core component of camera-based 3D understanding in autonomous driving, where view transformation (VT) maps multi-camera image features into a unified BEV representation. Sampling-based view transformation (Sampling-VT) is attractive because it supports dense and continuous BEV aggregation for high-resolution and long-range perception. Its deployment bottleneck, however, is systems-level: standard tensorized implementations of Sampling-VT -- which we refer to as Tensorized Sampling-VT -- explicitly materialize large height-dependent intermediate tensors, causing memory and latency costs that scale poorly with vertical resolution and the number of cameras. We revisit Tensorized Sampling-VT from an operator-execution perspective and show that it follows a gather-reduction pattern: each BEV query independently accumulates contributions across cameras and height bins, enabling thread-local accumulation with on-the-fly recomputation that eliminates the need to materialize height- and camera-dependent intermediates. Based on this insight, we propose FlashBEV, a fully fused and IO-aware execution strategy mathematically equivalent to Tensorized Sampling-VT (same operator output) while substantially reducing global memory traffic and kernel-launch overhead. Experiments show that FlashBEV achieves more than an order of magnitude lower peak GPU memory and significant inference-latency speedups, with memory effectively independent of the number of height bins, reducing the operator's peak memory to O(BCXY) (output only). This unlocks higher BEV range/resolution and vertical discretization within fixed deployment budgets on memory-constrained devices. Our contribution is an execution redesign -- same math, different execution -- that removes a key scalability barrier for deployment-ready Sampling-VT. Code available at https://github.com/yokosyun/FlashBEV

CommentsAccepted to ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑