StudyBench:自进化方法能否从教材中挖掘出奥林匹克竞赛级别的解题能力?
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
浏览论文内容
中文总结 AI 辅助
StudyBench是衡量自进化方法将教材转化为解题能力效率的物理基准,研究发现其存在指导差距与计算平台期,剩余差距源于方法本身,为自进化研究提供了可衡量目标。
中文摘要 AI 辅助
人类只需学习少量编写精良的教材就能掌握一门学科并尝试其最难的问题。我们认为,理想的自进化方法应具备相同特性,即能自主从原始训练材料中学习,形成可迁移的问题解决能力。然而,我们仍缺乏对此的直接衡量标准。我们推出StudyBench,这是一个受控的物理基准,直接衡量自进化方法将训练材料转化为能力的效率。我们将测试集分为应用集(由难度较高的教材问题组成,用于评估吸收能力)和迁移集(由奥林匹克竞赛级问题组成,用于评估迁移能力)。我们在三个基础模型上对代表性自进化方法进行基准测试,发现应用集上的提升很少能迁移到更难的迁移集上。一项指导消融实验揭示了指导差距:即使是最强的方法,在将相同材料作为上下文指导提供时,也仅能挖掘出其全部潜力的一小部分。此外,所有方法都达到了计算平台期,在耗尽计算预算前就已饱和。因此,剩余差距是方法问题,而非数据或计算问题。StudyBench通过提供一个干净且受控的基准,将自进化的进展从开放式探索转变为未来研究可衡量的目标。我们的代码已在此httpsURL发布。
英文摘要
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. We argue that an ideal self-evolution method should share the same property, that is autonomously learning from raw training material for transferable problem-solving capability. However, we still lack a direct measurement for it. We introduce StudyBench, a controlled physics benchmark that directly measures how efficiently a self-evolution method converts training material into capability. We organise the test set into an Application Set, consisting of difficult textbook problems and evaluating absorption ability, and a Transfer Set, consisting of olympiad-level problems and evaluating transfer ability. Benchmarking representative self-evolution methods across three base models, we find that improvements on the Application Set rarely translate to the harder Transfer Set. A guidance ablation exposes a Guidance Gap: even the strongest method closes only a small fraction of what the same material unlocks when supplied as in-context guidance. Besides, every method hits a Compute Plateau, saturating well before exhausting its compute budget. The remaining gap is therefore a method problem rather than a data or compute problem. By offering a clean and controlled benchmark, StudyBench turns self-evolution progress from an open-ended pursuit into a measurable target for future research. Our code is released at https://github.com/thunlp/StudyBench.
发表机构
- Tsinghua University(清华大学)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。