arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CASCADE:用于驾驶的时空因果推理表示与数据集

CASCADE: A Spatio-Temporal-Causal Reasoning Representation and Dataset for Driving

Jenny Schmalfuss, Despoina Paschalidou, Simon Gerstenecker, German Ros, Jose M. Alvarez

arXiv 2609.07094首次发表:更新:

发表机构

NVIDIA; ETH Zürich(英伟达; 苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CASCADE提出一种时空因果推理表示及人工标注数据集,用于驾驶场景,使推理预测可机器验证,并支持物理AI模型推理能力的基准测试。

AI 中文摘要

推理是自动驾驶在长尾场景中实现泛化的一条有前景的途径,因为它可以推断场景中的元素如何相互依赖,并遍历这些依赖关系以得出超出观察范围的结论。然而,很难判断模型的结论是否遵循场景的依赖关系,因为没有一种驾驶表示能将这些依赖关系明确到足以进行测试的程度。基于文本的推理轨迹缺乏时空基础,时空场景图缺乏因果联系,而大规模的推理标注越来越多地由模型生成且难以验证。为此,我们提出了CASCADE(驾驶环境的因果时空分析),它包含两个组成部分:(1)一种用于驾驶场景推理的结构化场景表示;(2)一个基于该表示构建的人工标注数据集。对于每个与自车交互的参与者,CASCADE表示会逐帧记录,只要该参与者可见,其采取的动作、发生的位置以及如何依赖于其他参与者的动作和状态。由此产生的结构使推理预测可被机器验证:可以逐元素地对其进行评分,而无需依赖(多模态)大语言模型评判。CASCADE数据集为PhysicalAI数据集的2,066个驾驶片段提供了全面的人工标注,包含超过34K个元素,建立了每个场景的时空和因果上下文,其中包括8.6K个带时间戳的自车和智能体动作、3.7K个因果链接和2.9K个潜在影响,以及6.1K个针对智能体、物体、交通灯和环境的标注。由于完全由人工标注,CASCADE为这一比较提供了参考:基准测试物理AI模型的推理能力,并验证自动生成的推理标签的质量。CASCADE数据集可在以下网址获取:此https URL。

英文摘要

Reasoning is a promising route to the generalization that autonomous driving requires in the long tail, as it can infer how the elements of a scene depend on one another and traverse those dependencies to conclusions beyond what is observed. Yet it is hard to tell whether a model's conclusions follow the scene's dependencies, because no driving representation makes them explicit enough to test against. Text-based reasoning traces lack spatio-temporal grounding, spatio-temporal scene graphs lack causal links, and reasoning annotations at scale are increasingly model-generated and hard to verify. To this end, we introduce CASCADE (Causal Spatio-Temporal Analysis of Driving Environments), which encompasses two components: (1) a structured scene representation for reasoning in driving scenes and (2) a human-annotated dataset built on it. For every actor that interacts with the ego vehicle, the CASCADE representation records frame-by-frame, for as long as the actor is visible, what action is taken, where it occurs, and how it depends on the actions and states of others. The resulting structure makes reasoning predictions machine-verifiable: they can be scored against it element by element, without relying on (M)LLM judges. The CASCADE dataset provides comprehensive human annotations for 2,066 driving clips of the PhysicalAI dataset, with over 34K elements that establish the spatio-temporal and causal context of each scene, including 8.6K time-stamped ego and agent actions, 3.7K causal links and 2.9K potential influences, and 6.1K annotations for agents, objects, traffic lights, and environments. Being entirely human-annotated, CASCADE provides the reference for this comparison: benchmarking the reasoning abilities of Physical AI models, and verifying the quality of automatically generated reasoning labels. The CASCADE dataset is available at https://huggingface.co/datasets/nvidia/cascade.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑