arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AIMO可解释性挑战

AIMO Interpretability Challenge

Michal Štefánik, Philipp Mondorf, Andreas Waldis, Qianying Liu, Chuan Yang, Michal Spiegel, Josef Kuchař, Marek Kadlčík, Adam Vawda-Oomerjee, Chaoran Liu, Simon Frieder, Barbara Plank, Fazl Barez, Pontus Stenetorp

arXiv 2607.13899首次发表:更新:

发表机构

National Institute of Informatics; Munich Center for Machine Learning / MaiNLP LMU; University of Tübingen; Fuzhou University; Masaryk University; University College London; University of Oxford(日本国立信息学研究所; 慕尼黑机器学习中心/慕尼黑大学语言与文学计算研究所; 图宾根大学; 福州大学; 马萨里克大学; 伦敦大学学院; 牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出AIMO可解释性挑战,基于前沿数学语言模型内部机制区分稳健与虚假推理。利用AIMO问题等资源,提供推理问题、模型访问及鲁棒性评估,助参与者开发识别稳健模型的方法,创建新基准和基线系统,连接可解释性与泛化研究。

AI 中文摘要

我们提出了AIMO可解释性挑战,这是一项基于前沿数学语言模型内部机制区分稳健推理和虚假推理的竞赛。该挑战源于标准推理基准的核心局限:高最终答案准确率无法揭示模型是依赖稳定推理机制还是利用脆弱推理捷径。基于人工智能数学奥林匹克(AIMO)问题及提交内容,结合菲尔兹模型计划的资源,竞赛将提供新发布的奥林匹克级数学推理问题及其符号表示、前沿推理模型访问权以及对模型在这些问题上的对抗鲁棒性评估。参与者将利用这些资源及计算基础设施支持,开发识别稳健解决问题模型的方法。竞赛还将创建新的开放鲁棒性基准和基线系统,旨在为数学推理和可解释性的标准基准测试提供持久基础。从科学角度看,竞赛围绕人工智能研究的核心问题连接了可解释性和泛化研究:我们能否确定前沿人工智能模型的决策在多大程度上是可泛化的,从而可靠?

英文摘要

We propose the AIMO Interpretability Challenge, a competition on distinguishing robust from spurious reasoning in frontier mathematical language models based on the models' internal mechanisms. The challenge is motivated by a central limitation of standard reasoning benchmarks: strong final-answer accuracy does not reveal whether a model relies on stable reasoning mechanisms or exploits brittle reasoning shortcuts. Building on AI Mathematical Olympiad (AIMO) problems and submissions, together with resources from the Fields Model Initiative, the competition will provide (1) newly-published olympiad-level math reasoning problems and their symbolic representations, allowing generation of novel functional variants, (2) access to frontier reasoning models, and (3) assessments of models' adversarial robustness on these problems. Participants will use these resources, along with our computing infrastructure support, to develop methods for identifying which models solve problems robustly. Our competition will also create a new, open robustness benchmark and baseline systems, aiming to provide a lasting foundation for standard benchmarking in mathematical reasoning and interpretability. Scientifically, the competition connects interpretability and generalization research around a central question in AI research: can we determine if, and to what extent, the decision-making of frontier AI models is generalizable and thus, reliable?

CommentsAccepted Competition at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑