arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25860cs.CVcs.RO

MatchFusion:时空多模态自动驾驶中的显式-隐式实例匹配

MatchFusion: Explicit-Implicit Instance Matching for Spatio-Temporal Multimodal Autonomous Driving

Xiaoyu Li, Jiajia Fu, Long Shi, Tianyu Du, Ruihang Li, Xian Wu, Lijun Zhao, Yingtao Zhang, Lining Sun, Ruifeng Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态自动驾驶中时空实例交互的匹配问题,提出MatchFusion模块,结合几何与语义的显式-隐式匹配,提升感知精度并大幅降低计算开销。

中文摘要 AI 辅助

稀疏实例表示为多模态感知和端到端自动驾驶(E2EAD)中的空间激光雷达-相机交互和时间过去-当前交互提供了紧凑的接口。有效的交互需要可靠的实例对应关系,尽管存在几何差异和异构语义表示。基于注意力的方法利用上下文语义,但通常需要专门的表示对齐,增加了计算开销。相比之下,基于结构化对象状态的关联高效且可解释,但缺乏解决模糊匹配的上下文证据。为了结合这些互补优势,我们提出了MatchFusion,一种用于时空多模态自动驾驶的可学习实例匹配与融合模块。MatchFusion利用几何相似性和类别一致性初始化成对亲和度,然后使用实例嵌入选择性地细化结构上合理的关联。生成的软匹配图指导一个通用的残差聚合算子进行自适应信息交换。这种统一的匹配-融合公式支持空间激光雷达-相机和时间过去-当前交互,分别使用多视图图像平面几何和运动补偿的鸟瞰图(BEV)几何作为结构先验。在nuScenes上的实验表明,在不同前端配置下均获得一致的感知提升。与先前的实例中心融合方法相比,配备MatchFusion的系统实现了更高的感知精度,同时将FLOPs降低了55.3%,GPU内存使用量降低了39.3%,匹配-融合模块仅占总感知延迟的3.7%。将时间MatchFusion集成到SparseDrive中,在端到端框架内进一步改善了感知,无需额外监督。这些结果确立了显式-隐式匹配作为时空实例交互的有效且高效的机制。

英文摘要

Sparse instance representations provide a compact interface for spatial LiDAR-camera and temporal past-current interaction in multimodal perception and E2EAD. Effective interaction requires reliable instance correspondences despite geometric discrepancies and heterogeneous semantic representations. Attention-based methods exploit contextual semantics but often require specialized representation alignment, increasing computational overhead. In contrast, association based on structured object states is efficient and interpretable but lacks contextual evidence to resolve ambiguous matches. To combine these complementary strengths, we propose MatchFusion, a learnable instance matching and fusion module for spatio-temporal multimodal autonomous driving. MatchFusion initializes pairwise affinities using geometric similarity and category consistency, then selectively refines structurally plausible associations using instance embeddings. The resulting soft matchmap guides a common residual aggregation operator for adaptive information exchange. This unified matching-fusion formulation supports spatial LiDAR-camera and temporal past-current interaction, using multi-view image-plane geometry and motion-compensated BEV geometry as the respective structural priors. Experiments on nuScenes demonstrate consistent perception gains across diverse front-end configurations. Compared with a prior instance-centric fusion method, the MatchFusion-equipped system achieves higher perception accuracy while reducing FLOPs by 55.3% and GPU memory usage by 39.3%, with the matching-fusion module accounting for only 3.7% of total perception latency. Integrating temporal MatchFusion into SparseDrive further improves perception within an E2E framework without additional supervision. These results establish explicit-implicit matching as an effective and efficient mechanism for spatio-temporal instance interaction.

发表机构

  • Harbin Institute of Technology, Harbin 150001, China(中国哈尔滨 哈尔滨工业大学 150001)
  • Zhejiang University, Hangzhou 310027, China(中国杭州 浙江大学 310027)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑