发表机构
DFKI; RPTU Kaiserslautern-Landau; BITS-Pilani, Hyderabad(德国人工智能研究中心; 凯泽斯劳滕-兰道工业大学; 比拉理工学院海得拉巴校区)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出NeuroSymbEAD,一个基于KITTI-360的大规模神经符号字幕数据集,通过自我中心知识图谱生成多层次场景字幕,为自动驾驶的视觉语言模型提供交通场景解释和3D推理基准。
AI 中文摘要
本文介绍了NeuroSymbEAD,一个大规模神经符号字幕数据集,其特点是一个以自我为中心的静态和动态对象知识图谱(KG),这些对象被标注了类别、分类、朝向、方向以及距自车(ego-vehicle)的距离。这些标注被应用于KITTI-360数据集,以生成多层次的文本字幕,代表了一种轻量级的以自我为中心的场景地图。户外场景地图重建、视觉识别和对象基础(object grounding)为驾驶常识和交通/场景理解建立了基线。为此,基于自然语言的对象及其复杂关系的基础字幕(grounded captioning)是室内场景任务中广泛采用的上下文表示方法。神经符号表示已被证明在处理各种计算机视觉和语言应用的结构化信息方面是有效的。我们的数据标注流程能够生成多样化的地图片段,在任何3D对象检测网络预测的边界框内填充模拟或真实对象,并构建层级文本字幕。我们使用预训练的基础(grounding)网络和学习到的自回归字幕网络对我们的神经符号和本体论字幕生成进行基准测试。通过将3D驾驶场景转换为结构化的以自我为中心的语言,NeuroSymbEAD为视觉语言和基础模型提供了交通场景解释、3D推理和可解释自动驾驶感知的基准。
英文摘要
This paper introduces NeuroSymbEAD, a large-scale neuro-symbolic caption dataset featuring an ego-centric knowledge graph (KG) of static and dynamic objects annotated with classes, categories, heading directions, orientations, and distances from the ego-vehicle. These annotations are used on the KITTI-360 dataset to generate multilevel textual captions representing a lightweight version of an ego-centric scene map. Outdoor scene-map reconstruction, visual recognition, and object grounding establish baselines for driving common sense and traffic/scene understanding. For these purposes, natural language-based grounded captioning of objects and their complex relationships is a widely adopted contextual representation for indoor scene tasks. Neuro-symbolic representations have proven effective in handling structured information for various computer vision and language applications. Our data annotation pipeline allows the generation of varied map segments, populating simulated or real objects within the bounding boxes predicted by any 3D object detection network, and building hierarchical text captions. We benchmark our neuro-symbolic and ontological caption generation using pre-trained grounding and learned auto-regressive captioning networks. By converting 3D driving scenes into structured ego-centric language, NeuroSymbEAD provides a benchmark for vision-language and foundation models for traffic-scene explanation, 3D reasoning, and interpretable autonomous-driving perception.