EventCoT:用于推理时间定位的以事件为中心的视频思维链
EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization
浏览论文内容
中文总结 AI 辅助
提出EventCoT框架解决推理时间定位难题,先对视频进行以事件为中心的分词,再在识别出的事件内推理生成答案,通过嵌入匹配确定时间间隔,在相关任务中取得好结果。
中文摘要 AI 辅助
推理时间定位(RTL)要求模型生成包含支持时间间隔的答案,需联合产生高级推理和精确时间定位。为此提出首个以事件为中心的视频思维链框架EventCoT,先对输入视频进行以事件为中心的分词,再在识别出的事件内推理生成答案,通过嵌入匹配确定时间间隔,在ActivityNet-RTL上取得领先结果,在ReXTime基准测试中也有出色的零样本结果。
英文摘要
Reasoning temporal localization (RTL) requires a model to generate an answer that itself contains the time interval supporting it, coupling high-level reasoning with temporal grounding in a single response. To tackle this challenge, we propose the first event-centric video chain-of-thought framework, dubbed EventCoT. EventCoT first performs event-centric tokenization, converting the video into compact event tokens that enable efficient identification of question-relevant events. It then reasons within these events to generate the answer, grounding the time interval via embedding matching that aligns placeholder tokens with visual embeddings. EventCoT achieves state-of-the-art results on ActivityNet-RTL while using substantially fewer visual tokens than previous work, and attains strong zero-shot results on the grounded video question answering benchmark ReXTime. Our code will be released for research purposes.
发表机构
- Pohang University of Science and Technology(浦项科技大学)
- Korea Advanced Institute of Science and Technology(韩国科学技术院)
- Handong Global University(韩东国际大学)
机构由 AI 辅助整理,请以论文原文为准。