arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TimePLE:重新思考视频时间定位的时间表示

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding

Yuhui Zeng, Xinyu Mao, Xiaokun Liu, Xin Tao, Jinfa Huang, Jiayi Ji, Xiawu Zheng

arXiv 2607.23951首次发表:更新:

AI 中文总结

研究视频时间定位问题,提出TimePLE方法,通过预测有效时间区间的联合分布将VTG从端点预测转为区间原生定位,经实验验证该方法优于端点预测基线,在短和中等持续时间事件上表现出色。

AI 中文摘要

视频时间定位(VTG)旨在定位自然语言查询描述的连续视频区间。当前基于VLM的方法通常通过两个端点输出间接生成此区间。我们提出TimePLE,它通过预测有效时间区间上的单个联合分布,将VTG从端点预测重新表述为区间原生定位。TimePLE将每个区间映射到规范位置 - 持续时间正方形中的一个点。给定视频和查询,VLM生成单个潜在<|TIMESPAN|>令牌,其隐藏状态解码为联合区间分布,通过持续时间感知坐标校正进行细化并转换为连续边界。使用相同区间表示对输入时间锚进行编码。为可靠对齐潜在跨度表示与完整事件区间,精心策划90K规模的有基础样本并人工验证3K规模的基准注释。实验表明TimePLE持续优于端点预测基线,在短持续时间和中等持续时间事件上有明显提升,平均mIoU达到58.9。

英文摘要

Video temporal grounding (VTG) aims to localize the continuous video interval described by a natural-language query. However, current VLM-based methods typically produce this interval indirectly through two endpoint outputs, represented either as discrete timestamp tokens or continuous boundary coordinates. These formulations differ in how endpoints are encoded, but not in what is predicted: the event interval remains a derived object, while interval validity, duration, and interval-level similarity are handled only implicitly. We propose TimePLE, which reformulates VTG from endpoint prediction to interval-native grounding by predicting a single joint distribution over valid temporal intervals. TimePLE maps each interval to a point in a canonical position-duration square, where every support point corresponds to a valid span and neighboring points represent geometrically similar intervals. Given a video and query, the VLM generates a single latent <|TIMESPAN|> token whose hidden state is decoded into a joint interval distribution, refined through duration-aware coordinate correction, and converted into continuous boundaries. The same interval representation is used to encode input temporal anchors, aligning video-side temporal evidence with output-side span prediction. To reliably align the latent span representation with complete event intervals, we curate 90K-scale grounded samples and human-verify 3K-scale benchmark annotations. Experiments across four VTG benchmarks show that TimePLE consistently outperforms endpoint prediction baselines, achieving an average mIoU of 58.9, with clear gains on short-duration and medium-duration events.

Comments25 pages, 13 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑