arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自由配方极限:每种配方效应都衡量理想化学习者的哪个前提被打破

The Free-Recipe Limit: Every Recipe Effect Measures Which Premise of an Idealised Learner Broke

Wenhui Chen, Jianlin Chen, Ziyao Lin, Chi Man Vong

arXiv 2609.26160首次发表:更新:

发表机构

University of Macau; South China University of Technology(澳门大学; 华南理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过大规模微调实验,测量配方搜索的极限,发现连贯性破坏是导致配方效应系统性变化的关键,而体积是唯一可靠提升性能的杠杆。

AI 中文摘要

固定一个语料库并将配方搜索推向无穷:尝试技能的每一种顺序、从块状到交错的各种排列、每一种组合,并保留最佳结果。两个量决定了该搜索的价值:其探索的可达集合的直径,以及任何人区分两个端点所需的分辨率。当直径低于分辨率时,再多的搜索也无法转化为决策,其标志并非没有赢家,而是赢家无法在重新运行中存活。我们通过12个基础模型(0.5B-14B,三个预训练家族)上的761次微调运行来测量这一配方搜索之墙,涉及竞赛数学技能:基础检查点、AdamW下的监督微调、k=4时的精确匹配评分。在固定体积的单一连贯领域内,三个经典自由度平均为0.010-0.021,而基线为0.019;最大的对比度为0.0619,超过了三种子分辨率,并在两次重新运行中分别读出+0.010和-0.015。具有系统性答案的偏离是连贯性:将合并的语料库减半,并让两半在不兼容但同样正确的约定下撰写答案,将排列从能力转移到约定之间的分配,比同约定对照高两个数量级,而将约定写入输入则关闭了这一现象。该切换在第二个预训练家族上重复,并在独立重新执行其自身协议后仍然成立,重新执行散布(0.087)小于搜索选择的顺序单元携带的分辨率(0.144)。顺序本身是一个瞬态,其符号在单次运行内三次过零。体积,这个无人称之为配方的杠杆,是唯一可靠回报的杠杆。一个公开记分卡对所有26项预注册声明进行评级:18项支持,5项失败,2项未测试,1项混合。

英文摘要

Fix a corpus and send recipe search to infinity: try every order of the skills, every arrangement from blocked to interleaved, every composition, and keep the best. Two quantities decide what that search was worth: the diameter of the reachable set it explores, and the resolution at which anyone can tell two endpoints apart. Where the diameter falls below the resolution, no amount of search converts into a decision, and the signature is not an absence of winners but winners that do not survive re-running. We measure this recipe-search wall with 761 fine-tuning runs on 12 base models (0.5B-14B, three pretraining families) over competition-mathematics skills: base checkpoints, supervised fine-tuning under AdamW, exact-match scoring at k=4. Within one coherent domain at fixed volume the three classical freedoms average 0.010-0.021 against a 0.019 floor, and the largest contrast, 0.0619, clears a three-seed resolution and then reads +0.010 and -0.015 on two reruns. The departure with a systematic answer is coherence: halving one pooled corpus and letting the halves write answers under incompatible but equally correct conventions moves arrangement from capability to allocation between conventions, by two orders of magnitude over a same-convention control, and writing the convention into the input switches the phenomenon off. The switch replicates on a second pretraining family and survives an independent re-execution of its own protocol, with a re-execution spread (0.087) smaller than the resolution a search-selected order cell carries (0.144). Order itself is a transient whose sign crosses zero three times inside a single run. Volume, the one lever nobody calls a recipe, is the one that reliably pays. A public scorecard grades all 26 pre-registered claims: 18 supported, 5 failed, 2 untested, 1 mixed.

Comments42 pages, 9 figures, 17 tables (incl. appendices); code, pre-registration files, and result base to be released with the paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑