arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

我在视频中寻找你:面向以人为中心的视频推理的身份条件查询

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

Shibo Gao, Chongxiao Wang, Chenglong Huang, Jie Ma, Haolin Shi, Fei Ding, Jing Li, Qiang Lyu, Yangyang Liu, Yang Liu, Jun Liu, Linlin Huang, Peipei Yang

arXiv 2608.07417首次发表:更新:

发表机构

Beijing Jiaotong University; HUJING Digital Media & Entertainment Group; MAIS Institute of Automation, Chinese Academy of Sciences(北京交通大学; 汇晶数字媒体与娱乐集团; 中国科学院自动化研究所MAIS)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出了身份条件查询(ICQ)任务,构建了ISYV基准、训练集与框架,实验显示主流多模态大语言模型在该任务上表现不佳,ISYV模型性能优于强基线且接近闭源模型。

AI 中文摘要

现实世界的视频推理通常涉及多模态、多源输入,而现有的视频推理任务通常采用简化的视频-文本设置,限制了身份匹配和以人为中心的推理。为了弥合这一差距,我们提出了身份条件查询(Identity-conditioned Queries,ICQ)任务,要求模型联合关联并解释输入视频和某个人的参考图像,利用该条件解决身份 grounding、行为理解、时间推理等挑战。基于ICQ,我们提出了ISYV(I Seek You in Videos),这是一个包含三个组件的系统性解决方案:(1)ISYV-Bench,一个具有挑战性的评估基准,包含1377个真实世界复杂视频和1377个问答对,分为六个难度级别,涵盖从身份识别到因果推理的能力;(2)ISYV-75K,一个大规模训练集,包含75K高质量样本,通过自动标注、多阶段验证和人工审核构建;(3)ISYV-Framework,包含面向ICQ的模型和训练策略,用于学习利用有信息的视频片段,无需额外的片段级标注。大量实验表明,主流的闭源和开源多模态大语言模型(MLLM)在ISYV-Bench上表现不佳,尤其是在跨域身份匹配和长时跟踪方面。ISYV-Model的性能优于强基线,在某些方面接近闭源模型的性能。总体而言,ISYV为以人为中心的视频推理提供了统一的任务定义、可扩展的数据集/基准以及建模见解。

英文摘要

Real-world video reasoning often involves multimodal, multi-source inputs, whereas existing video reasoning tasks typically assume a simplified video-text setting, limiting identity matching and person-centric reasoning. To bridge this gap, we introduce the Identity-conditioned Queries (ICQ) task, in which models are required to jointly associate and interpret an input video and a reference image of a person, and leverage this conditioning to address identity grounding, behavior understanding, and temporal reasoning, among other challenges. Building on ICQ, we present ISYV (I Seek You in Videos), a systematic solution comprising three components: (1) ISYV-Bench, a challenging evaluation benchmark with 1,377 real-world complex videos and 1,377 question-answer pairs, organized into six difficulty levels spanning capabilities from identity recognition to causal reasoning; (2) ISYV-75K, a large-scale training set of 75K high-quality samples constructed via automated annotation, multi-stage verification, and manual review; and (3) ISYV-Framework, containing an ICQ-oriented model and training strategy for learning to exploit informative video shots without additional shot-level annotations. Extensive experiments show that both mainstream closed-source and open-source MLLMs struggle on ISYV-Bench, especially in cross-domain identity matching and long-horizon tracking. ISYV-Model outperforms strong baselines and in some aspects approaches closed-source performance. Overall, ISYV provides a unified task definition, scalable datasets/benchmarks, and modeling insights for person-centric video reasoning.

CommentsAccepted to ACM Multimedia 2026 (MM '26). 6 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑