发表机构
INRIA Paris(法国国家信息与自动化研究所巴黎分部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究用问答式问题测试多语言大语言模型在各类主题上的表现,引入TriviaRoomQA基准评估其日常等知识,发现模型在知识密集型主题强,流行文化主题弱,且性能因语言而异,揭示了现有基准未捕捉的知识差距。
AI 中文摘要
问答室、知识竞赛之夜和问答节目挑战着人类在从经典事实到日常文化等广泛主题上的知识。本文使用问答式问题在常见和小众主题上测试大语言模型(LLMs),以检验它们在这类场景中能否具有竞争力。我们引入了TriviaRoomQA,这是一个多语言基准,旨在评估288个主题的日常、基于文化和长尾知识。该基准包含六种欧洲语言的3300个平行多项选择题以及另外5340个仅法语的问题用于更细致的案例研究。我们评估了来自欧洲、亚洲和北美的30个开放权重的LLMs,涵盖7到70B参数的模型。我们发现模型在历史、地理和数学等知识密集型主题上表现强劲,但在名人、音乐、电影和新闻等日常流行文化主题上则弱得多。此外,即使对于相同的基础问题,模型性能也因语言而异,这表明获取事实知识并不总是与语言无关。总之,我们的数据集和实验证明了现有基于学术的饱和基准未捕捉到的重要知识差距。
英文摘要
Quiz rooms, trivia nights, and quiz shows challenge human knowledge across a wide range of topics, from canonical facts to everyday culture. In this paper, we examine whether large language models (LLMs) can perform competitively in such settings, using quiz-style questions to test them on both common and niche topics. We introduce TriviaRoomQA, a multilingual benchmark designed to evaluate everyday, culturally grounded, and long-tail knowledge across 288 topics. The benchmark contains 3,300 parallel multiple-choice questions in six European languages and additional 5,340 French-only questions for a more fine-grained case study. We evaluate 30 open-weight LLMs from European, Asian, and North American providers, covering models from 7 to 70B parameters. We find that models are strong on knowledge-intensive topics such as history, geography, and mathematics, but substantially weaker on everyday popular-culture topics such as celebrities, music, movies, and news. Moreover, model performance varies across languages even for the same underlying questions, suggesting that access to factual knowledge is not always language-independent. In sum, our dataset and experiments demonstrate an important knowledge gap which is not captured by existing academic-based saturated benchmarks.
CommentsEMNLP 2026 Findings