arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01948cs.CV

Event ActivityNet:用于未修剪动作理解的大规模模拟事件基准

Event ActivityNet: A Large-Scale Simulated-Event Benchmark for Untrimmed Action Understanding

Cheng-Yao Hong, Ting-Wei Lin, Yun-Chung Lai, Hua-Wei Lee, Hwann-Tzong Chen, Tyng-Luh Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出大规模模拟事件基准Event ActivityNet,解决长时序事件动作理解数据集不足的问题,建立基线并验证其预训练对原生事件微调的性能提升效果。

中文摘要 AI 辅助

长时序基于事件的动作理解研究仍未得到充分探索,原因在于现有数据集大多由短的、经过修剪的片段组成,而收集带有密集时间标注的原生事件流成本高昂。我们推出Event ActivityNet,这是一个源自人工标注的未修剪ActivityNet视频的大规模模拟事件基准。它包含3263个视频、200个动作类别,总时长106.94小时,配有匹配的5-bin和9-bin事件体素表示、时间动作标注以及带时间戳的字幕。该基准支持标注片段动作识别、辅助事件-语言对齐以及因果在线时间动作定位。我们按解码帧顺序直接从非插值源视频生成事件体素,保留每个视频合理的标称或平均帧率元数据用于近似时间映射,并使用动作中心重建LPIPS作为保留可重建内容的软诊断指标。我们为自适应事件 framing、提示-字幕对齐以及仅事件、仅RGB和RGB-事件定位建立了基线。在渐进嵌套尺度训练协议下,识别Top-1准确率从52.25提升至66.42,在线时间定位平均mAP从21.7提升至29.0。此外,分阶段Event ActivityNet预训练后再进行原生事件微调,在多种监督预算下始终优于仅目标训练和从头联合训练。Event ActivityNet为长时序事件建模提供了可扩展的基准,不过面向部署的结论仍需原生相机评估。

英文摘要

Long-horizon event-based action understanding remains underexplored because existing datasets largely comprise short, trimmed clips, while collecting native event streams with dense temporal annotations is costly. We introduce Event ActivityNet, a large-scale simulated-event benchmark derived from human-annotated, untrimmed ActivityNet videos. It comprises 3,263 videos, 200 action classes, and 106.94 hours, with matched 5-bin and 9-bin event-voxel representations, temporal action annotations, and timestamped captions. The benchmark supports annotated-segment action recognition, auxiliary event-language alignment, and causal online temporal action localization. We generate event voxels directly from non-interpolated source videos in decoded frame order, retain per-video rational nominal or average frame-rate metadata for approximate time mapping, and use action-center reconstruction LPIPS as a soft diagnostic of retained reconstructable content. We establish baselines for adaptive event framing, prompt-caption alignment, and event-only, RGB-only, and RGB-event localization. Under a progressive nested-scale training protocol, recognition Top-1 accuracy increases from 52.25 to 66.42, while online temporal localization average mAP improves from 21.7 to 29.0. Moreover, staged Event ActivityNet pretraining followed by native-event fine-tuning consistently outperforms target-only and joint-from-scratch training across multiple supervision budgets. Event ActivityNet provides a scalable benchmark for long-horizon event modeling, although native-camera evaluation remains essential for deployment-oriented conclusions.

补充信息

↑