Video-MME-Logical: 面向视频时序逻辑推理的受控诊断基准
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning
- HKUST(香港科技大学)
- Colab, Beihang University(北京航空航天大学合作实验室)
- CUHK(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
为隔离评估多模态大模型在视频中的时序逻辑推理能力,提出包含五种时序逻辑操作和25个细粒度任务类别的受控基准,通过控制时间跨度和推理复杂度进行难度可控评估,并支持中间状态诊断。
AI中文摘要:
近期对多模态大语言模型(MLLMs)的兴趣引发了一个核心问题:它们能否对动态视觉证据进行推理,而不仅仅是识别单帧中的物体或事件?这种能力,我们称之为视频时序逻辑推理,要求模型在视觉状态随帧演变时维护、更新和组合证据。现有的视频基准常常将这种能力与场景复杂度、静态识别或不受控的时间变化混为一谈。为了隔离这种能力,我们引入了Video-MME-Logical,一个围绕五种时序逻辑操作组织的受控基准:状态跟踪、顺序计数、时间排序、动态空间性和结构组合。该基准包含25个细粒度任务类别,通过受控的对象状态、转换、时间依赖和逻辑组合生成。它通过变化时间跨度和推理复杂度实现难度可控的最终答案评估,并通过验证模型在产生最终答案前是否恢复了所需的逻辑推理轨迹来支持中间状态诊断。与最先进的MLLMs的实验揭示了显著的人机差距,尤其是当时序逻辑复杂度增加时。在多达50万个生成样本上的监督微调提高了性能,但仍不足以弥合理赔差距,这使得Video-MME-Logical成为分析和改进MLLMs中时序逻辑推理的可扩展测试平台。
英文摘要:
Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely recognize objects or events in individual frames? This ability, which we refer to as video temporal-logical reasoning, requires models to maintain, update, and compose evidence as visual states evolve across frames. Existing video benchmarks often conflate this capability with scene complexity, static recognition, or uncontrolled temporal variation. To isolate this capability, we introduce Video-MME-Logical, a controlled benchmark organized around five temporal-logical operations: state tracking, sequential counting, temporal ordering, dynamic spatiality, and structural composition. The benchmark contains 25 fine-grained task categories generated with controlled object states, transitions, temporal dependencies, and logical compositions. It enables difficulty-controlled final-answer evaluation by varying temporal horizon and reasoning complexity, and supports intermediate-state diagnostics by verifying whether models recover the required logical reasoning trace before producing the final answer. Experiments with state-of-the-art MLLMs reveal a substantial human-model gap, especially as temporal-logical complexity increases. Supervised fine-tuning on up to 500K generated samples improves performance but remains insufficient to close the reasoning gap, positioning Video-MME-Logical as a scalable testbed for analyzing and improving temporal-logical reasoning in MLLMs.