大语言模型通过了历史考试却遗漏了《历史》:波兰高中毕业考试Matura基准测试
Lost in Historical Time? A Polish History Matura Benchmark for Large Language Models
浏览论文内容
中文总结 AI 辅助
该研究针对波兰Matura历史考试评估8款LLM,发现其总分远超人类考生,但存在波兰历史得分惩罚、源内容混淆等缺陷,推出了首个基于波兰国家课程的LLM历史基准。
中文摘要 AI 辅助
AI聊天机器人被学生广泛用作知识来源,但大语言模型(LLM)基准测试很少评估解释性历史推理。我们在波兰高中毕业考试Matura历史科目上评估了8款领先的LLM,涵盖2023-2025年的3份官方试卷,包括简答题和长篇作文,将模型表现与人类考生群体进行对比。所有模型的表现都远超人类考生,但总分掩盖了不同的能力特征:在任务类型、源模态和地理范围方面的排名不稳定,且在波兰历史内容上始终存在得分惩罚,对比世界历史内容表现更差。定性错误分析揭示了两种反复出现的失败模式:源内容混淆,即模型从源内容进行推理,而非将其作为分析对象;时间定位错误,即回答存在历史错位。本研究推出了首个基于波兰国家课程的LLM历史基准测试。
英文摘要
Language models are widely used by students as knowledge sources, yet benchmarks rarely assess their interpretative historical reasoning. We evaluate eight leading LLMs on the Polish high school exit exam (Matura) in history - three official papers from 2023-2025, comprising short-answer questions and extended essays - and compare model performance against the human examinee population. Although models score near the ceiling, aggregate scores mask distinct competency profiles: rankings are unstable across task types, source modalities, and geographical scopes, with a consistent penalty for Polish versus Global history content. Qualitative error analysis reveals two recurring failure modes - source decontextualization, when models reason from source content rather than treating it as an object of analysis, and temporal disorientation, when responses are historically misplaced. This study introduces the first LLM history benchmark grounded in the Polish national curriculum.
发表机构
- Adam Mickiewicz University(亚当·密茨凯维奇大学)
- IDEAS Research Institute(IDEAS研究院)
- AMU Center for Artificial Intelligence(亚当·密茨凯维奇大学人工智能中心)
- Poznań University of Medical Sciences(波兹南医科大学)
- Maria Curie-Skłodowska University(玛丽·居里-斯克洛多夫斯卡大学)
- WSB Merito University(WSB梅里托大学)
机构由 AI 辅助整理,请以论文原文为准。