arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

模态成熟度指数:评估全模态模型多模态能力的基准

Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models

Rohit Patel, Dieuwke Hupkes, Sloan Strader

arXiv 2608.26317首次发表:更新:

发表机构

Meta Superintelligence Labs(元超级智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出评估全模态模型多模态能力的基准MMI,测试五种模态及组合,发现前沿模型MPS得分低,LLM评分者与人工标注者判断一致性达70.8%。

AI 中文摘要

前沿语言模型越来越多地被宣传为可跨模态感知和响应的全模态系统。然而,现有的评估框架几乎只关注双模态理解,通常是文本加另一种模态。我们提出了模态成熟度指数(Modality Maturity Index,MMI),这是一个旨在评估大型语言模型多模态能力的基准,涵盖五种模态(文本、图像、音频、视频和文档),以及输入和输出中最多三种模态的组合。MMI包含893个问题,每个问题都经过精心设计,要求模型展示对多种输入模态的理解,并生成包含各种输出格式的响应。这些问题设计为自包含,对于准确响应所需的正确模态或模态组合有明确的预期。每个MMI提示都带有针对响应中预期的每种输出模态的人工编写的评分标准;模型的MMI值是每个提示的各模态得分的平均值。由于低分可能反映无法生成某一模态(存在性不足)或无法生成正确内容,我们还引入了补充性的模态存在性得分(Modality Presence Score,MPS),即每个提示针对预期输出模态的F1值。将MMI应用于五个前沿多模态模型,我们发现MPS的范围仅为15.6(Claude Opus 4.6)至34.9(GPT-5.4)。考虑到可用于评分的返回模态可用性较低,我们报告MPS为主要结果,等待模型改进。为了评估用LLM评分者和评分标准判断输出正确性的可行性,我们使用自定义生成工具进行了单独实验。在生成的样本中,我们发现应用评分标准的LLM评分者与不使用评分标准的人工标注者(直接对输出评分且从未查看评分标准)的判断一致性为70.8%。

英文摘要

Frontier language models are increasingly marketed as omni systems that can perceive and respond across modalities. Existing evaluation frameworks, however, focus almost exclusively on bimodal understanding, typically text plus one other modality. We propose the Modality Maturity Index (MMI), a benchmark designed to evaluate the multimodal capabilities of large language models across five modalities (text, image, audio, video and document) and combinations of up to three modalities in both inputs and outputs. MMI consists of 893 questions, each carefully crafted to require the model to demonstrate its understanding of multiple input modalities and to generate responses that incorporate various output formats. The questions are designed to be self-contained, with clear expectations for the correct modality or mix of modalities required for an accurate response. Every MMI prompt carries human-authored rubric criteria for each output modality expected in the response; a model's MMI Value expresses the average of the per-modality scores for each prompt. Because low scores can reflect either failure to generate a modality (lack of presence) or failure to generate correct content, we introduce also a supplementary Modality Presence Score (MPS), a per-prompt F1 over the expected output modalities. Applying MMI to five frontier multimodal models, we find that the MPS ranges from only 15.6 (Claude Opus 4.6) to 34.9 (GPT-5.4). Given the low availability of returned modalities to even grade, we report MPS as our main result pending model improvements. To assess the viability of judging output correctness with LLM judges and rubrics, we run a separate experiment with custom generation tools. On the assets that generates, we find that an LLM judge applying the rubrics agrees with rubric-blind human annotators (who score the outputs directly and never see the criteria) on 70.8% of judgments.

Comments26 pages, 6 figures. Code and dataset available

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑