ProactiveBench:流式视频模型真的能像人类一样交互吗?
LiveProBench: Can Streaming Video Models Really Interact Like Humans?
- Beihang University(北京航空航天大学)
- Alibaba(阿里巴巴)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对流式视频模型缺乏主动交互评估的问题,提出ProactiveBench基准,以一秒间隔无提示评估模型,发现多数系统过早响应多于遗漏,揭示时间决策差距。
AI中文摘要:
流式视频理解要求模型在处理连续多模态输入的同时保持时间上下文。现有评估主要是反应式的:它们在选定的时间戳查询模型,因此不评估模型何时应该响应。主动交互则要求模型监控持续请求,在目标事件发生后的适当时间间隔内响应,否则保持沉默。我们提出了ProactiveBench,它在没有明确响应提示的情况下,以一秒的流间隔评估模型。其六个子任务在触发模糊性和时间容忍度上有所不同。事件敏感性在同一记录上几何组合了响应率和沉默率;四个基于窗口的子任务区分了早期、窗口内和遗漏响应;重复计数则惩罚遗漏和重复。在六个被评估的系统中,有四个系统的过早响应数量超过了遗漏响应数量,揭示了在类人交互所需的时间决策方面存在显著差距。
英文摘要:
Streaming video understanding requires models to process continuous multimodal input while maintaining temporal context. Existing evaluations are predominantly reactive: they query a model at a selected timestamp and therefore do not assess when it should respond. Proactive interaction instead requires monitoring a standing request, responding within an appropriate interval after the target event, and otherwise remaining silent. We introduce LiveProBench, which evaluates models at one-second stream intervals without an explicit response cue. Its six subtasks vary trigger ambiguity and timing tolerance. Event Sensitivity geometrically combines response and silence rates on the same recording; four window-based subtasks distinguish early, in-window, and missed responses; and Duplicate Counting penalizes omissions and repetitions. Premature responses outnumber missed responses for half of the evaluated models, revealing a substantial gap in the temporal decision-making required for human-like interaction.