arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OVIBench:面向中断场景的在线视频问答基准测试

OVIBench: Benchmarking Online Video Question Answering under Interruption

Naiming Liu, Zhiheng Wu, Shuning Wang, Tie Zhang, Bowen Liu, Tong Wang

arXiv 2608.22279首次发表:更新:

发表机构

HIT; CASIA; ZJU; UESTC; HKUST(哈尔滨工业大学; 中国科学院自动化研究所; 浙江大学; 电子科技大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有视频问答基准未考虑用户中断的问题,提出首个面向中断场景的在线视频问答基准OVIBench,开发模拟协议与指标集,构建训练集OVI-Train,验证了其有效性。

AI 中文摘要

近期,视觉语言模型(VLMs)在视频理解领域取得了显著进展。然而,现有的大多数视频问答研究及基准测试仍遵循离线、单轮范式,忽略了用户可能在模型生成答案过程中进行中断的真实交互场景。为填补这一空白,本文提出了面向中断场景的在线视频问答任务,并引入OVIBench——首个用于评估该场景下VLMs的标准化基准测试。OVIBench将中断分为三类:取消(Cancellation)、误触发(False Trigger)、修正(Correction),同时支持开放式问答和多项选择问答两种评估形式。为实现大规模且可复现的测试,本文开发了一种离线模拟协议,该协议在统一时间设置下复现模型生成过程中的中断,同时配套了多维度指标集,用于评估模型的中断理解能力与响应生成质量。实验表明,OVIBench能够有效区分模型的中断处理能力,尤其是在遵循修正请求方面。最后,本文构建了用于中断感知微调的训练集OVI-Train,在该数据集上微调后的模型在OVIBench上取得了显著提升,验证了基准测试与数据设计的有效性。OVIBench、OVI-Train及评估代码将被开源发布。

英文摘要

Recent vision language models (VLMs) have achieved strong progress in video understanding. However, most existing video QA research and benchmarks still follow an offline, single-round paradigm, overlooking realistic interactions where users may interrupt the model during answer generation. To address this gap, we formulate the task of Online Video Question Answering under Interruption and introduce OVIBench, the first standardized benchmark for evaluating VLMs in this setting. OVIBench categorizes interruptions into three types: Cancellation, False Trigger, Correction and supports both open-ended and multiple-choice evaluations. To enable large-scale and reproducible testing, we develop an offline simulation protocol that reproduces interruption during generation under a unified temporal setup, together with a multi-dimensional metric suite for assessing interruption understanding and response generation. Experiments demonstrate that OVIBench effectively distinguishes models' interruption-handling abilities, especially in following correction requests. Finally, we construct a train set OVI-Train for interruption-aware fine-tuning. Models fine-tuned on this dataset achieve significant gains on OVIBench, validating the effectiveness of our benchmark and data design. OVIBench, OVI-Train, and the evaluation code will be released.

CommentsEMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑