arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19380cs.CV

CAViAR:用于真实场景细粒度事故推理的因果视频数据集

CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios

Sparsh Garg, Yi-Wen Chen, Vijay Kumar B G, Abhishek Aich

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出人工标注的真实事故视频基准CAViAR,测试发现现有VLMs存在感知-推理差距,无法可靠将驾驶场景主体行为映射到责任类别。

中文摘要 AI 辅助

现代自动驾驶系统在目标检测、轨迹预测等感知任务中表现出色,但缺乏解释交通事故所需的高级因果推理能力。特别是确定责任归属,如识别谁有过错、违反了哪条交通规则,在当前基准中仍未得到充分探索。为此,我们推出CAViAR(因果事故视频与事件分析库,Causal Accident Video and Incident Analysis Repository),这是一个人工标注的行车记录仪基准,包含从CarCrashDataset(CCD)和Nexar收集的2249个真实事故视频。每个视频都标注了结构化标签,涵盖环境条件、事故类型、因果解释、明显过错主体、受影响主体以及明显违规类别。我们对最先进的视觉语言模型(VLMs)进行基准测试,包括Cosmos-Reason2、Qwen3-VL和InternVL3。一旦通过多数类/随机基线和平衡指标解决类别不平衡问题,感知能力表现不均——光照问题几乎已解决,而天气和道路状况的准确率等于或低于多数类基线——所有模型在事故类型和责任推理上的性能都急剧下降。总体而言,CAViAR揭示了实际的感知-推理差距:当前的VLMs可能能识别显著的上下文,但在安全关键的驾驶场景中,无法可靠地将可见的主体行为映射到带标注的与规则相关的责任类别。代码、标注方案、提示和评估脚本可在以下网址获取:this https URL

英文摘要

While modern autonomous driving systems excel at perception tasks such as object detection and trajectory prediction, they lack the high-level causal reasoning required to interpret traffic accidents. In particular, determining responsibility, such as identifying who is at fault and which traffic rule was violated, remains largely unexplored in current benchmarks. To this end, we introduce CAViAR (Causal Accident Video and Incident Analysis Repository), a human-annotated dashcam benchmark comprising 2,249 real-world accident videos collected from CarCrashDataset (CCD) and Nexar. Each video is annotated with structured labels spanning environmental conditions, accident type, causal explanation, apparent At-Fault Agent, affected agent, and apparent rule-violation category. We benchmark state-of-the-art vision-language models (VLMs), including Cosmos-Reason2, Qwen3-VL, and InternVL3. Once class imbalance is accounted for with majority/random baselines and balanced metrics, perceptual competence is uneven--lighting is nearly solved, whereas weather and road-condition accuracy fall at or below the majority-class baseline---and all models degrade sharply on accident type and responsibility reasoning. Overall, CAViAR exposes a practical Perception--Reasoning Gap: current VLMs may recognize salient context, but do not reliably map visible agent actions to annotated rule-relevant responsibility categories in safety-critical driving scenarios. Code, annotation schema, prompts, and evaluation scripts are available at: https://github.com/nec-labs-ma/CAViAR

发表机构

  • NEC Laboratories, America(美国NEC实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑