发表机构
Beijing Jiaotong University; HUJING Digital Media & Entertainment Group; MAIS Institute of Automation, Chinese Academy of Sciences(北京交通大学; 汇晶数字媒体与娱乐集团; 中国科学院自动化研究所MAIS)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出了身份条件查询(ICQ)任务,构建了ISYV基准、训练集与框架,实验显示主流多模态大语言模型在该任务上表现不佳,ISYV模型性能优于强基线且接近闭源模型。
AI 中文摘要
现实世界的视频推理通常涉及多模态、多源输入,而现有的视频推理任务通常采用简化的视频-文本设置,限制了身份匹配和以人为中心的推理。为了弥合这一差距,我们提出了身份条件查询(Identity-conditioned Queries,ICQ)任务,要求模型联合关联并解释输入视频和某个人的参考图像,利用该条件解决身份 grounding、行为理解、时间推理等挑战。基于ICQ,我们提出了ISYV(I Seek You in Videos),这是一个包含三个组件的系统性解决方案:(1)ISYV-Bench,一个具有挑战性的评估基准,包含1377个真实世界复杂视频和1377个问答对,分为六个难度级别,涵盖从身份识别到因果推理的能力;(2)ISYV-75K,一个大规模训练集,包含75K高质量样本,通过自动标注、多阶段验证和人工审核构建;(3)ISYV-Framework,包含面向ICQ的模型和训练策略,用于学习利用有信息的视频片段,无需额外的片段级标注。大量实验表明,主流的闭源和开源多模态大语言模型(MLLM)在ISYV-Bench上表现不佳,尤其是在跨域身份匹配和长时跟踪方面。ISYV-Model的性能优于强基线,在某些方面接近闭源模型的性能。总体而言,ISYV为以人为中心的视频推理提供了统一的任务定义、可扩展的数据集/基准以及建模见解。
英文摘要
Real-world video reasoning often involves multimodal, multi-source inputs, whereas existing video reasoning tasks typically assume a simplified video-text setting, limiting identity matching and person-centric reasoning. To bridge this gap, we introduce the Identity-conditioned Queries (ICQ) task, in which models are required to jointly associate and interpret an input video and a reference image of a person, and leverage this conditioning to address identity grounding, behavior understanding, and temporal reasoning, among other challenges. Building on ICQ, we present ISYV (I Seek You in Videos), a systematic solution comprising three components: (1) ISYV-Bench, a challenging evaluation benchmark with 1,377 real-world complex videos and 1,377 question-answer pairs, organized into six difficulty levels spanning capabilities from identity recognition to causal reasoning; (2) ISYV-75K, a large-scale training set of 75K high-quality samples constructed via automated annotation, multi-stage verification, and manual review; and (3) ISYV-Framework, containing an ICQ-oriented model and training strategy for learning to exploit informative video shots without additional shot-level annotations. Extensive experiments show that both mainstream closed-source and open-source MLLMs struggle on ISYV-Bench, especially in cross-domain identity matching and long-horizon tracking. ISYV-Model outperforms strong baselines and in some aspects approaches closed-source performance. Overall, ISYV provides a unified task definition, scalable datasets/benchmarks, and modeling insights for person-centric video reasoning.
CommentsAccepted to ACM Multimedia 2026 (MM '26). 6 figures, 5 tables