arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoRE:驾驶视频中的弱监督由粗到细风险证据学习

CoRE: Weakly Supervised Coarse-to-Fine Risk Evidence Learning in Driving Videos

Kaiser Hamid, Can Cui, Nade Liang

arXiv 2608.25344首次发表:更新:

发表机构

Texas Tech University; Purdue University(德克萨斯理工大学; 普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CoRE是一种弱监督由粗到细框架,可从粗粒度视频监督中学习驾驶视频的细粒度风险证据,在三个基准测试中表现出色,无需耗时的细粒度标注即可实现风险定位。

AI 中文摘要

驾驶过程中感知风险会随时间演变,且可能由特定场景实体支撑,但监督信号通常仅局限于粗粒度的视频级判断。学习支撑证据出现的时机以及哪些实体支撑风险预测器,通常需要耗时的时间级和实体级标注。我们提出了**CoRE**,一种弱监督由粗到细框架,可从粗粒度视频监督中学习细粒度预测支撑。CoRE首先训练一个视频级预测器,随后将其冻结;对候选时间区域或实体轨迹进行结构化干预,测量每个候选对象如何改变粗粒度预测,从而产生分级的预测效果目标;将这些目标蒸馏到学生模型中,使其能直接从原始视频预测时间和实体支撑,无需在推理阶段进行干预。我们在三个互补设置中评估该学习原理:RISEE测试从主观片段级判断中获取感知风险支撑,无时间或实体级风险标注;DoTA提供独立的时间事件标注,用于评估弱监督交通异常定位;UCF-Crime测试同一由粗到细机制是否可扩展至标准非驾驶异常检测基准。在这些设置中,CoRE从粗粒度监督中学习到了有用的细粒度支撑,在DoTA上实现了出色的时间定位,在UCF-Crime上取得了有竞争力的性能。这些结果表明,粗粒度视频预测可提供有用的监督,用于恢复支撑它们的细粒度证据,无需相应的细粒度标签。

英文摘要

Perceived risk in driving evolves over time and may be supported by specific scene entities, yet supervision is typically limited to coarse video-level judgments. Learning \emph{when} supporting evidence emerges and \emph{which entities} support a risk predictor would ordinarily require costly temporal- and entity-level annotations. We introduce \textbf{CoRE}, a weakly supervised coarse-to-fine framework that learns fine-grained prediction support from coarse video supervision. CoRE first trains a video-level predictor and then freezes it. Structured interventions over candidate temporal regions or entity tracks measure how each candidate changes the coarse prediction, producing graded prediction-effect targets. These targets are distilled into a student that directly predicts temporal and entity support from the original video, without requiring interventions at inference. We evaluate this learning principle across three complementary settings: RISEE tests perceived-risk support from subjective clip-level judgments without temporal or entity-level risk annotations; DoTA provides independent temporal event annotations for evaluating weakly supervised traffic-anomaly localization; and UCF-Crime tests whether the same coarse-to-fine mechanism extends to a standard non-driving anomaly-detection benchmark. Across these settings, CoRE learns informative fine-grained support from coarse supervision, with strong temporal localization on DoTA and competitive performance on UCF-Crime. These results show that coarse video predictions can provide useful supervision for recovering the fine-grained evidence supporting them, without requiring corresponding fine-grained labels.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑