AI 中文总结
提出基于EventGraph和EventField的结构化时间视频推理流水线,在EPIC-KITCHENS子集上准确率达0.98,优于字幕基线和VLM,兼顾性能与可解释性。
AI 中文摘要
我们提出了一种结构化的时间视频推理流水线,该流水线围绕离散的 EventGraph、连续的 EventField 以及人类可读的 EventGlyph 视图构建。在由 10 个视频和 50 个时间推理问题组成的校准后的 EPIC-KITCHENS 子集上,EventField+Glyph 实现了 0.98 的总体准确率,比字幕基线高出 +0.40(配对 p = 1.1 × 10^{-5}),比直接仅用 VLM 的问答高出 +0.20(p = 0.0063)。我们进一步评估了标注来源的变体,包括手动、启发式和启发式+Gemini 流水线,发现最佳的结构化方法在所有设置下均保持在字幕基线之上。我们还进行了跨视频对基准测试,并提供了所有研究视频的 glyph 输出附录图库。总体而言,结果表明,通过保留符号结构、捕获时间连续性并提供人类可读的视频推理诊断,结构化时间表示能够同时支持性能和可检查性。
英文摘要
We present a structured temporal video reasoning pipeline built around a discrete EventGraph, a continuous EventField, and a human-readable EventGlyph view. On a calibrated EPIC-KITCHENS subset of 10 videos and 50 temporal reasoning questions, EventField+Glyph achieves 0.98 overall accuracy, which is higher than the caption baseline by +0.40 (paired p = 1.1 \times 10^{-5}) and direct VLM-only QA by +0.20 (p = 0.0063) on this subset. We further evaluate annotation-source variations, including manual, heuristic, and heuristic+Gemini pipelines, and find that the best structured method stays above the caption baseline across settings. We also include cross-video pair benchmarking and an appendix gallery of glyph outputs for all studied videos. Overall, the results indicate that structured temporal representations can support both performance and inspectability by preserving symbolic structure, capturing temporal continuity, and providing human-readable diagnostics for video reasoning.
CommentsSubmitted to WACV 2027. Preprint; 9 pages, 7 figures. Copyright may be transferred without notice