arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于自动驾驶的多模态场景相似性搜索

Multimodal Scenario Similarity Search for Autonomous Driving

Tamás Matuszka, András Tamásy, Balázs Szolár

arXiv 2607.09428首次发表:更新:

发表机构

aiMotive(爱莫提夫)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究自动驾驶场景相似性搜索,提出多模态框架结合视觉与轨迹表示,对比两种基于轨迹的方法与视觉基线,实验表明轨迹表示在运动事件中性能强,视觉嵌入在外观线索丰富时出色,结合二者可提升检索质量。

AI 中文摘要

大规模自动驾驶数据集包含大量记录的场景,需要高效检索方法来识别与给定查询相似的情况。现有方法通常依赖视觉表示或基于运动的描述,难以理解其在场景检索中的相对优缺点。本文提出一个多模态框架,在统一检索管道中结合视觉和基于轨迹的表示。研究了两种基于轨迹的方法:Exo-Trajectory(基于周围代理运动的显式匹配方法)和ScenarioFormer(基于对比学习从对象轨迹学习的基于Transformer的表示)。将这些方法与强大的基于视觉的基线进行比较,并在各种驾驶场景中分析其行为。实验结果表明,轨迹表示在切入、转弯操作和交通排队等以运动为中心的事件中提供强大的检索性能,而视觉嵌入在外观线索信息丰富时表现出色。最重要的是,结合视觉和轨迹信息持续提高检索质量,产生最佳整体性能。这些发现表明外观和运动捕捉是场景相似性的互补概念,并推动了用于自动驾驶数据挖掘、数据集整理和基于场景验证的多模态检索系统。

英文摘要

Large-scale autonomous-driving datasets contain vast numbers of recorded scenarios, creating a need for efficient retrieval methods that can identify situations similar to a given query. Existing approaches typically rely on either visual representations or motion-based descriptions, making it difficult to understand their relative strengths and limitations for scenario retrieval. In this work, we present a multimodal framework for autonomous-driving scenario retrieval that combines visual and trajectory-based representations within a unified retrieval pipeline. We investigate two trajectory-based approaches: Exo-Trajectory, an explicit matching method based on surrounding-agent motion, and ScenarioFormer, a transformer-based representation learned from object trajectories using contrastive learning. We compare these approaches against strong vision-based baselines and analyze their behavior across a diverse set of driving scenarios. Experimental results show that trajectory representations provide strong retrieval performance for motion-centric events such as cut-ins, turning maneuvers, and traffic queueing, while visual embeddings excel when appearance cues are informative. Most importantly, combining visual and trajectory information consistently improves retrieval quality, yielding the best overall performance. These findings demonstrate that appearance and motion capture are complementary notions of scenario similarity and motivate multimodal retrieval systems for autonomous-driving data mining, dataset curation, and scenario-based validation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑