发表机构
Indian Institute of Technology Jodhpur(印度焦特布尔印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SynCrash多阶段流水线,针对CVPR2026挑战赛ACCIDENT任务,在无真实标注数据下,结合VideoMAEv2、YOLO及物理启发式算法,实现交通监控视频零样本事故检测、定位与碰撞分类。
AI 中文摘要
我们提出SynCrash,一种用于固定视角CCTV监控视频中零样本事故检测、空间定位及碰撞类型分类的多阶段流水线。该方法针对CVPR 2026挑战赛ACCIDENT任务,要求在无标注真实训练数据的情况下,预测事故发生时间、帧内碰撞位置及碰撞类型。流水线分为三个解耦阶段:(1)基于VideoMAEv2-giant骨干网络,在基于CARLA生成的合成片段上微调,结合元数据感知嵌入与稠密滑动窗口推理实现时间定位;(2)使用YOLO目标检测结合物理信息混合启发式算法,利用边界框重叠与轨迹推理预测碰撞点,实现空间定位;(3)基于检测到的车辆数量与配置,采用轻量规则策略进行碰撞类型分类。核心见解是:时间理解受益于合成数据上的监督微调,而空间理解更适合使用跨域可自然迁移的预训练目标检测器与物理先验。
英文摘要
We present SynCrash, a multi-stage pipeline for zero-shot accident detection, spatial localization, and collision-type classification in fixed-view CCTV surveillance video. Our approach addresses the ACCIDENT at CVPR 2026 Challenge, which requires predicting when an accident occurs, where in the frame the impact happens, and what type of collision it is, all without access to labeled real-world training data. The pipeline operates in three decoupled stages: (1) Temporal localization via a VideoMAEv2-giant backbone fine-tuned on CARLA-based synthetic clips with metadata-aware embeddings and dense sliding-window inference; (2) Spatial localization using YOLO for object detection combined with a physics-informed hybrid heuristic that leverages bounding-box overlap and trajectory-based reasoning to predict the impact point; and (3) Collision-type classification using a lightweight rule-based strategy derived from the number and configuration of detected vehicles. The key insight is that temporal understanding benefits from supervised fine-tuning on synthetic data, whereas spatial understanding is better served by pretrained object detectors and physics priors that transfer naturally across domains.
CommentsAccepted at the CVPR 2026 AUTOPILOT Workshop (non-archival)