arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Kairos:面向空间、时间与动态的细粒度视频-语言建模数据集

Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics

Ruibo Ming, Lei Sun, Deheng Zhang, He Zhang, Jialu Li, Jian Wang, Zhendong Li, Mengshun Hu, Danda Pani Paudel, Luc Van Gool, Jinjin Gu

arXiv 2609.08755首次发表:更新:

发表机构

INSAIT; Sofia University “St. Kliment Ohridski”; Adobe Research; Snap Research(INSAIT; 索菲亚大学“圣·克利门特·奥赫里德斯基”; 奥多比研究院; Snap研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Kairos数据集通过细粒度时间对齐标注的长视频,支持持续动作、实体与交互建模,为视频语言理解提供通用基础。

AI 中文摘要

许多新兴的视频语言建模任务要求系统超越片段级抽象,并在扩展的时间范围内对视觉内容进行建模。然而,大多数现有的视频数据集依赖于粗略或稀疏对齐的监督,这压缩了时间变化,并限制了模型学习连续视觉动态的可重用表示的能力。我们引入了Kairos,一个具有时间分辨标注的视频数据集,用于视频语言建模。Kairos包含时长从十分钟到半小时的长视频,并标注了细粒度的时间对齐。这些标注捕获了沿视频时间线的持续动作、实体出现和属性、交互以及不断演变的上下文线索。这种时间分辨结构支持细粒度评估、长程建模和推理、指令数据构建、表示学习和视频生成。Kairos为随时间建模视觉体验提供了通用基础。

英文摘要

Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most existing video datasets rely on coarse or sparsely aligned supervision, which compresses temporal variation and limits the ability of models to learn reusable representations of continuous visual dynamics. We introduce Kairos, a video dataset for video-language modeling with time-resolved annotations. Kairos consists of long-duration videos, ranging from ten minutes to half an hour, annotated with fine-grained temporal alignment. The annotations capture ongoing actions, entity appearances and attributes, interactions, and evolving contextual cues along the video timeline. This time-resolved structure supports fine-grained evaluation, long-range modeling and reasoning, instruction data construction, representation learning, and video generation. Kairos provides a general-purpose foundation for modeling visual experiences over time.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑