BEAR-Bench:面向多模态模型的双语企业与学术推理基准
BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models
浏览论文内容
中文总结 AI 辅助
本文针对多模态大语言模型在文本密集专业文档推理能力评估的不足,推出英俄双语的BEAR-Bench基准,评估16个多模态大语言模型并对比幻觉检测方法,发现最强模型仍有明显性能提升空间。
中文摘要 AI 辅助
尽管多模态大语言模型(MLLMs)在视觉理解方面已取得显著进展,但它们对文本密集型专业文档的推理能力仍未得到充分评估。现有基准侧重信息提取、依赖外部领域知识,或仅将专业文档作为众多场景之一,且大多以英语或中文为中心,其他语言尤其是俄语的代表性严重不足。为解决这些局限,本文推出BEAR-Bench(Bilingual Enterprise and Academic Reasoning,双语企业与学术推理基准),这是一个自成体系的复杂英俄双语基准,包含1000个人工标注问题,基于文本丰富的商业及科学文档构建。我们在BEAR-Bench上评估了16个专有及开源权重的多模态大语言模型,包括Gemini 3.1 Pro和Qwen3.5-397B,即便对于最强的系统,也观察到明显的性能提升空间。最后,我们利用所得模型输出对比现有幻觉检测方法,不仅评估模型在BEAR-Bench上的失败频率,还评估识别这些失败的可靠性。
英文摘要
While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings. They are also largely English- or Chinese-centric, leaving other languages and Russian, in particular, substantially underrepresented. To address these limitations, we introduce BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents. We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, on BEAR-Bench and observe clear headroom even for the strongest systems. Finally, we use the resulting model outputs to compare existing hallucination detection methods, evaluating not only how often models fail on BEAR-Bench but also how reliably those failures can be identified.
发表机构
- Yandex Applied AI Institute(Yandex应用人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。