发表机构
Hunyuan, Tencent; Nanyang Technological University; Northumbria University(腾讯混元; 南洋理工大学; 诺森比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对MLLM视听推理评估的痛点,提出OmniCapBench深度结构化评估框架,通过原子评估单元实现可靠评分,揭示前沿MLLM局部感知强但长时视听推理弱的问题,为全模态发展提供路线图。
AI 中文摘要
多模态大语言模型(MLLM)正快速向连续视听推理方向发展,亟需能暴露其能力极限的评估方法。视听字幕生成是理想的诊断任务,但现有基准存在耦合权衡:整体字幕评分有覆盖度却无定位,局部探针有定位却无覆盖度,无约束的LLM评判则存在不稳定性。我们提出OmniCapBench(全视频字幕基准),该基准将视听字幕评估重构为深度结构化诊断框架,将预测目标从自由文本转换为三个轨迹上的原子、可验证评估单元集合:实体引用、视觉镜头和音频事件,通过确定性约束检查与基于LLM的语义比较实现可靠评分。OmniCapBench包含786个密集标注视频,可有效区分MLLM的感知错误,包括时间定位失败、身份漂移、跨模态错位和幻觉描述。对前沿MLLM的评估显示其局部感知能力强,但长时视听推理能力弱,尤其在身份漂移和跨模态错位方面,为全模态发展提供了细粒度路线图。
英文摘要
Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide coverage without localization, local probes provide localization without coverage, and unconstrained LLM judges introduce instability. We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic framework. OmniCapBench shifts the prediction target from free-form text to sets of atomic, verifiable evaluation units across three tracks: entity references, visual shots, and audio events, enabling reliable scoring with deterministic constraint checks and localized LLM-based semantic comparisons. With 786 densely annotated videos, OmniCapBench effectively distinguishes MLLM perception errors, including temporal grounding failures, identity drift, cross-modal misalignment, and hallucinated descriptions. Evaluating frontier MLLMs reveals strong local perception but weak long-horizon audio--visual reasoning, particularly in identity drift and cross-modal misalignment, providing a fine-grained roadmap for omnimodal development.
CommentsAccepted by NeurIPS 2026. Code and benchmark can be found at https://01yzzyu.github.io/OmniCapBench/