arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Video-IFBench:评估多模态大语言模型在视频理解场景中的指令遵循能力

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

Hongbo Liu, Peixian Chen, Sihan Liu, Peiyuan Zhang, Kai Zou, Dian Zheng, Xiaoxing Hu, Yuhao Dong, Mengdan Zhang, Yunhang Shen, Haoyu Cao, Wei Liu, Weibo Gu, Xing Sun, Shengjie Zhao

arXiv 2608.25529首次发表:更新:

发表机构

TJU; Tencent Youtu Lab; SJTU; Tencent Hunyuan; CUHK; NTU(天津大学; 腾讯优图实验室; 上海交通大学; 腾讯混元; 香港中文大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文推出Video-IFBench基准,构建含四类模板的指令分类体系与半自动数据 pipeline,评估20余种MLLMs的视频指令遵循能力,发现当前模型在复杂指令上仍具挑战,推动相关研究。

AI 中文摘要

多模态大语言模型(MLLMs)在视频理解领域已展现出较强性能,但它们在该领域的指令遵循能力仍未得到充分探索。现实场景中的视频理解要求模型不仅能正确解读视频内容,还需满足用户指定的各类约束条件。现有基准主要聚焦于任务准确率,而非指令遵循情况,导致该能力未得到充分评估。为填补这一空白,本文推出Video-IFBench——一个用于评估视频理解场景中指令遵循能力的综合基准,要求模型必须满足用户指定的各类约束,包括基于视觉和音频内容的约束。本文构建了包含四种模板的指令分类体系,涵盖单任务、多任务、选择型及嵌套指令,涉及32种任务类型和39类人工设计的约束类别,覆盖语义与格式要求。为降低标注成本,本文搭建了结合MLLMs、程序化处理及人工验证的半自动数据构建流程,最终生成1500个样本。本文对20余种近期MLLMs开展大规模评估,结果表明视频指令遵循对当前模型仍具挑战性,尤其在包含大量约束、语义约束或需基于视频内容选择正确分支/路径的复杂条件结构指令上。本文希望该工作能推动未来视频理解场景中指令遵循相关研究的发展。

英文摘要

Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not only to interpret video content correctly, but also to satisfy diverse user-specified constraints. Existing benchmarks focus primarily on task accuracy rather than instruction adherence, leaving this capability insufficiently evaluated. To address this gap, we introduce Video-IFBench, a comprehensive benchmark for evaluating instruction following in video understanding, where models must satisfy diverse user-specified constraints, including those grounded in visual and audio content. We develop an instruction taxonomy with four templates, including single-task, multi-task, selection, and nested instructions, covering 32 task types and 39 manually designed constraint categories spanning both semantic and format requirements. To reduce annotation cost, we build a semi-automatic data construction pipeline that combines MLLMs, programmatic processing, and human verification, resulting in 1.5K samples. We conduct a large-scale evaluation of more than 20 recent MLLMs and show that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content. We hope our work will facilitate future research on instruction following in video understanding scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑