arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ArchitectureIQ:论训练直觉的度量

ArchitectureIQ: On the Measure of Training Intuition

Zirui Ren, Shaoyang Guo, Chencheng Tang, Jinxin Wang, Chengyu Xiong, Shanbin Yu, Peihang Li, Yidi Wu, Bangzhe Huang, Qingyu Qu, Leqian Yang, Ziming Liu

arXiv 2609.39714首次发表:更新:

发表机构

MetaCircle; Tsinghua University; Peking University; University of California, Berkeley; Fudan University(MetaCircle; 清华大学; 北京大学; 加州大学伯克利分校; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ArchitectureIQ基准以度量LLM与人类的模型训练直觉,发现LLM直觉良好但存在不完美、经验性、非压缩及对数据不敏感四大局限,并构建知识库可显著提升弱模型性能。

AI 中文摘要

顶尖研究人员拥有良好的直觉,但语言模型在模型训练方面是否拥有与顶尖AI研究人员一样好的直觉?为了度量LLM和人类的模型直觉,我们引入了ArchitectureIQ基准。每个问题呈现一个合成数据集和若干训练方案,测试者被要求预测能产生最佳测试指标的训练方案。总体而言,我们发现LLM的模型直觉良好但存在四个局限:(1)直觉不完美,在某些情况下甚至低于人类。前沿模型达到约76%的准确率(随机选择为33%),而最佳人类研究者为66.0%,但仍远未达到完美。对于仅涉及架构的问题,最佳人类达到65%,而GPT-6 Astra仅有38%。(2)直觉是经验性的而非结构化的,这由更多CoT计算并未带来显著改进这一事实所支持。与数学不同,我们仍缺乏一种“AI科学”语言来实现对AI的结构化推理。(3)直觉并非最大程度地压缩,可进一步压缩为知识库。我们构建的仅含20个条目的知识库为弱模型带来了巨大收益:配备累积知识的GPT-4o几乎达到了Claude Opus 5的性能。(4)直觉对数据集属性不敏感,但最佳模型通常应依赖于数据属性。这表明数据是AI中真正的“暗物质”——LLM(人类研究人员同样如此)对数据的理解太少,甚至不如对模型架构的理解。

英文摘要

Top researchers have good intuition, but do language models have as good intuition about model training as top AI researchers? To measure model intuition of LLMs and humans, we introduce the ArchitectureIQ benchmark. Each question presents a synthetic dataset and several training recipes, and the test-taker is asked to predict the recipe yielding the best test metric. Overall, we find that LLMs' model intuition is good but has four limitations: (1) The intuition is imperfect, or even sub-human in some cases. Frontier models achieve around 76% accuracy (random choice 33%) vs best human researcher (66.0%), yet remain far from perfect. For architecture-only questions, best human achieves 65% while GPT-6 Astra only has 38%. (2) The intuition is empirical, not structured, supported by the fact that more CoT compute does not lead to substantial improvement. Unlike math, we still lack a "Science of AI" language that enables structured reasoning on AI. (3) The intuition is not maximally condensed, and can be further compressed into a knoledge base. Our constructed knowledge base with only 20 items yields large gains for weak models: GPT-4o equipped with the accumulated knowledge almost matches the performance of Claude Opus 5. (4) The intuition is insensitive to dataset properties, but the best model should in general depend on data properties. This suggests that data is the real "dark matter" in AI -- LLMs (so do human researchers) understand too little about data, even less than model architectures.

Comments29 pages, 10 figures. Code and reproduction materials: https://github.com/renrua52/ArchitectureIQ

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑