dKFD:面向固定预算局部事件理解的相位结构化证据分配
dKFD: Phase-Structured Evidence Allocation for Fixed-Budget Localized Event Understanding
浏览论文内容
中文总结 AI 辅助
针对固定预算下局部事件视频的稀疏证据选择,提出相位结构化可微选择器dKFD,通过前事件、事件和后事件相位分配证据,显著提升帧AUC并保持识别与定位性能。
中文摘要 AI 辅助
稀疏视频理解通常需要在固定帧预算下选择一小部分视觉证据。大多数稀疏选择器全局分配这一预算,允许所有帧相互竞争。对于时间局部化事件,这可能是一种较差的归纳偏置:有用的证据通常分布在前事件上下文、事件本身和后事件后果中。我们研究了局部化事件视频的固定预算证据分配,并表明全局竞争的Top-K选择器在保留识别和定位能力的同时,会产生不稳定的事件证据。在DoTA视频异常识别上,全局Top-K在K=12时获得了有竞争力的识别和时间定位,但选择器-事件对齐度较低(帧AUC为51.5±10.8)。我们提出了dKFD,一种相位结构化的可微选择器,在全序列时间编码后,在前事件、事件和后事件相位之间保留证据容量。在匹配预算的多种子评估中,dKFD在K=12时比匹配的全局Top-K选择器将帧AUC提高了+30.97(p<0.01),同时产生了适度但统计显著的识别增益和可比较的时间定位。机制消融表明相位监督是承重墙:即使保留相位划分预算,移除它也会将帧AUC降低到41.1±12.0。在VRU-Accident上的下游诊断显示,跨VLM家族相比学习的全局Top-K有一致的增益,而密集字幕揭示了一个边界条件,其中均匀采样仍然具有竞争力。这些结果支持相位结构化分配作为事件中心稀疏证据选择的受控固定预算方法,而非通用的视频摘要策略。
英文摘要
Sparse video understanding often requires selecting a small set of visual evidence under a fixed frame budget. Most sparse selectors allocate this budget globally, allowing all frames to compete with one another. For temporally localized events, this can be a poor inductive bias: useful evidence is often distributed across pre-event context, the event itself, and post-event consequences. We study fixed-budget evidence allocation for localized event videos and show that globally competitive Top-$K$ selectors can preserve recognition and grounding while producing unstable event evidence. On DoTA Video Anomaly Recognition, Global Top-$K$ obtains competitive recognition and temporal grounding, but low selector-event alignment at $K=12$ (Frame AUC $51.5 \pm 10.8$). We propose dKFD, a phase-structured differentiable selector that reserves evidence capacity across pre-event, event, and post-event phases after full-sequence temporal encoding. Under matched-budget multi-seed evaluation, dKFD improves Frame AUC by $+30.97$ over a matched Global Top-$K$ selector at $K=12$ ($p<0.01$), while yielding modest but statistically significant recognition gains and comparable temporal grounding. Mechanism ablations show that phase supervision is load-bearing: removing it reduces Frame AUC to $41.1 \pm 12.0$ even when phase-partitioned budgets are retained. Downstream diagnostics on VRU-Accident show consistent gains over learned Global Top-$K$ across VLM families, while dense captioning reveals a boundary condition where uniform sampling remains competitive. These results support phase-structured allocation as a controlled fixed-budget approach for event-centric sparse evidence selection, not as a universal video summarization strategy.