CLFTv2:基于层次化特征金字塔的高效相机-激光雷达语义分割融合
CLFTv2: Efficient Camera-LiDAR Fusion for Semantic Segmentation via Hierarchical Feature Pyramids
- Tallinn University of Technology(塔林理工大学)
- FinEst Centre for Smart Cities, Tallinn University of Technology(塔林理工大学芬埃斯特智慧城市中心)
- Universitas Mercatorum(梅尔卡托鲁姆大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
CLFTv2提出层次化相机-激光雷达融合框架,用Swin多尺度编码器和轻量FPN解码器替代全局注意力,在三个数据集上提升VRU召回率,计算量更低、吞吐量更高,适用于实时车载感知。
AI中文摘要:
自动驾驶中的语义分割需要在严重的类别不平衡情况下可靠地检测弱势道路使用者(VRU)。我们提出了CLFTv2,一种层次化的相机-激光雷达融合框架,用基于Swin的多尺度编码器和轻量级FPN风格残差解码器替代了全局ViT注意力。CLFTv2在2D透视域中运行,通过移位窗口注意力和逐尺度残差融合集成多尺度几何线索,避免了查询匹配解码器的计算开销。在三个驾驶数据集上,CLFTv2持续提高了VRU召回率。在ZOD上,CLFTv2-Large达到了53.5%的mIoU,将行人IoU从之前的CLFT模型的35.5%提高到44.9%。在Waymo上,CLFTv2达到了61.7%的mIoU。此外,一项模态隔离研究表明,ViT的全局感受野仅在密集激光雷达返回下产生更强的融合增益。与基于Swin的Mask2Former改编相比,CLFTv2所需的GFLOPs减少了1.4倍,吞吐量提高了2.2倍,同时实现了相当的整体精度。这些结果表明,层次化局部注意力融合为智能交通系统中的实时车载感知提供了一种高效、可扩展的替代全局注意力和基于查询的解码器的方案。源代码已公开。
英文摘要:
Semantic segmentation for autonomous driving requires reliable detection of vulnerable road users (VRUs) despite heavy class imbalance. We introduce CLFTv2, a hierarchical camera-LiDAR fusion framework replacing global ViT attention with a Swin-based multi-scale encoder and a lightweight FPN-style residual decoder. Operating in the 2D perspective domain, CLFTv2 integrates multi-scale geometric cues through shifted-window attention and per-scale residual fusion, avoiding the computational overhead of query-matching decoders. Across three driving datasets, CLFTv2 consistently improves VRU recall. On ZOD, CLFTv2-Large achieves 53.5\% mIoU, improving pedestrian IoU from 35.5\% to 44.9\% over the prior CLFT model. On Waymo, CLFTv2 reaches 61.7\% mIoU. Additionally, a modality-isolation study suggests ViT's global receptive field yields stronger fusion gains only under dense LiDAR returns. Compared to a Swin-based Mask2Former adaptation, CLFTv2 requires 1.4$\times$ fewer GFLOPs and delivers 2.2$\times$ higher throughput, while achieving comparable overall accuracy. These results demonstrate that hierarchical local-attention fusion offers an efficient, scalable alternative to global-attention and query-based decoders for real-time on-vehicle perception in intelligent transportation systems. Source code is publicly available.