OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs
OVO-S-Bench:多模态大语言模型中流式空间智能的分层基准
机构 * Tsinghua University(清华大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Beihang University(北京航空航天大学)
专题命中 仿真评测 :autonomous driving(abstract);分类 cs.CV
AI总结 提出OVO-S-Bench,一个完全人工标注的流式空间智能基准,包含1680个问题,涵盖四个抽象层次,评估38个MLLM,发现Gemini-3.1-Pro落后人类专家27分,流式空间微调MLLM表现不如其骨干模型。
Comments Accepted to EMNLP 2026 Main Conference. 55 pages, 12 figures, 20 tables. Project page: this https URL (https://internlm.github.io/OVO-S-Bench/)