arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03677cs.CVcs.CLcs.LGcs.NEcs.RO

通过自然语言描述图像子集间的差异来理解自动驾驶数据集

Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language

Julian Truetsch, Felix Hauser, Christoph Stiller, Frank Bieder

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出集差异描述任务,基于两阶段框架适配自动驾驶场景,引入基准 AD-Diff Bench,用开放权重模型开展实验,为自动驾驶数据集提供可解释的自省方法。

中文摘要 AI 辅助

理解大规模自动驾驶数据集的组成对于跨领域的安全性、鲁棒性和可靠运行至关重要。例如,不同地点间的域偏移可能导致运行环境与训练数据不匹配,进而引发潜在危险的性能下降。然而,现有的数据分析流程大多依赖元数据、预定义标签或人工检查,这些方式提供的语义洞察有限,且无法扩展。本文研究集差异描述任务:给定两个图像子集,目标是生成描述目标集与参考集之间差异的自然语言假设。基于两阶段框架,我们通过聚焦于目标检测得到的以对象为中心的图像块,将该方法适配到自动驾驶场景,这简化了聚合过程,并能将差异归因于特定的对象实例或类别。为了在域内评估该设置,我们引入了新的基准 AD-Diff Bench。低浓度实验评估了集差异描述方法对稀疏的真实世界差异的适用性。我们将实验限制在开放权重模型上,以支持可复现性和部署便捷性。所提出的基准和分析为自动驾驶数据集的实用、可解释的数据集自省提供了一步进展。我们的实现和基准数据集可在此 https URL 获取。

英文摘要

Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness, and reliable operation across domains. For example, domain shift between locations could lead to the operating environment being misaligned with the training data, resulting in potentially dangerous performance degradation. Yet, existing data analysis pipelines largely rely on metadata, predefined labels, or manual inspection, which provide limited semantic insight or do not scale. This paper studies set difference captioning: given two subsets of images, the goal is to produce a natural-language hypothesis describing differences between the target and reference set. Building on a two-stage formulation, we adapt the method to autonomous driving by focusing on object-centric patches derived from object detection, which simplifies aggregation and enables attribution of differences to specific object instances or categories. To evaluate this setting in-domain, we introduce a new benchmark, AD-Diff Bench. Low-concentration experiments assess the suitability of set-difference-captioning approaches to sparse, real-world differences. We restrict our experiments to open-weight models to support reproducibility and ease of deployment. The proposed benchmark and analysis provide a step towards practical, human-interpretable dataset introspection for autonomous driving datasets. Our implementation and benchmark dataset are available at https://github.com/KIT-MRT/AD-Diff

发表机构

  • FZI Research Center for Information Technology(FZI信息技术研究中心)
  • Karlsruhe Institute of Technology (KIT)(卡尔斯鲁厄理工学院(KIT))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑