IOL-AI挑战赛:一项旨在推进语言推理的开放挑战赛
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning
- University College London(伦敦大学学院)
- Meta
- CentraleSupélec(中央高等电力学院)
- McGill University(麦吉尔大学)
- Cohere Labs(Cohere实验室)
- Princeton University(普林斯顿大学)
- University of Amsterdam(阿姆斯特丹大学)
- Stockholm University(斯德哥尔摩大学)
- Faculty of Liberal Arts and Sciences(文理学院)
- Universidade Federal de Goiás(戈亚斯联邦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文介绍基于2026年IOL个人赛未公开题目的IOL-AI挑战赛,评估含自动与评审团评分,测试显示模型能力不由规模决定,语言推理是泛化推理技能的强基准代理。
AI中文摘要:
大型语言模型(LLM)的推理研究绝大多数集中在为模型提供规则的领域,如数学和代码。而语言谜题则与此相反:求解者必须先发现规则体系,再在其中进行推理。本文介绍了IOL-AI挑战赛,这是一项开放科学竞赛,基于2026年国际语言学奥林匹克(IOL)个人赛的未公开题目举办,采用自动评估与官方IOL评审团成员评估相结合的方式,且首次使用与人类参赛者相同的评分标准进行评审。该挑战赛在严格的计算预算(1块T4显卡,30分钟)下,共收到来自46支团队的731份提交作品。我们还对15个无资源限制的前沿模型和开放模型进行了基准测试,其中Claude Opus 4.8获得了与金牌相当的评审团分数,而我们提交给评审团评分的两个受资源限制的系统得分处于参赛者后5%的范围内。能力并非由模型规模决定:14B规模的提交作品的表现优于规模是其两倍的模型,性能提升来自解码和输出处理而非模型容量。我们还发现,自动指标对系统的排名与评审团的排名完全一致,但压缩了评分尺度,对弱系统的评分高出约13分,对强系统的评分则偏低。分析显示,尽管前沿模型可能对部分题目涉及的语言存在先验知识,但这并未显著帮助它们解决语言推理任务,这使得语言推理成为可泛化推理技能的强有力基准代理。
英文摘要:
Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We present the IOL-AI Challenge, an open-science competition run on the unseen problems of the International Linguistics Olympiad (IOL) 2026 Individual Contest, evaluated both automatically and, for the first time, by members of the official IOL Jury under the same rubrics applied to human contestants. The challenge drew 731 submissions from 46 teams under a strict compute budget (one T4, 30 mins). We additionally benchmark 15 unconstrained frontier and open models, with Claude Opus 4.8 earning a jury score equivalent to a gold medal, while both resource-constrained systems we submitted for jury grading scored in the range of the bottom 5% of contestants. Capability was not determined by scale: 14B submissions outperform models twice their size, and gains come from decoding and output-handling rather than model capacity. We also found that automatic metrics rank systems exactly as the jury does, but compress the scale, upscoring weak systems by ~13 points and understating strong ones. Our analysis shows that while frontier models might have prior knowledge about some of the problem languages, it does not significantly help them solve the linguistic reasoning tasks, leaving linguistic reasoning as a strong benchmarking proxy for generalizable reasoning skills.