arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29636cs.CLcs.AIcs.LGcs.SE

办公规模验证搜索中的操作符包、提出者强度与构造家族平台期

Operator Packages, Proposer Strength, and Construction-Family Plateaus in Office-Scale Verified Search

Roberto I. Ono Filho

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在办公规模验证搜索中,通过因子实验发现操作符组合(记忆+排斥)能显著缩小差距并防止崩溃,且前沿提出者优势显著,但家庭提示测试揭示构造家族平台期,循环可传递优化想法但无法自主产生。

中文摘要 AI 辅助

验证搜索(verified search)中,语言模型提出程序,硬性评估器对其进行评分,选择机制保留最优者,该方法近期已刷新数学纪录;然而,对提出方(proposer)组件的受控消融研究仍然稀少。我们在办公规模下(笔记本电脑上的30B本地模型,每次运行120-600个验证样本)实现了一个最小化的FunSearch风格循环,并配备了三个操作符包:模型编写并携带的示意性笔记本(而非逐字精英)、命名障碍(named obstacle),以及针对已发现构造的行为排斥(behavioural repulsion)。在来自公共仓库的九个构造问题上,完整的2^3因子设计(含两次重复)在名义上的两阶段分析中支持主要对比:该组合缩小了从初始种子到纪录的差距(+0.196;名义合并p=0.023,阶段组合p~0.08;每个问题的中位效应为+0.045)。排斥在所有地方提高了构造哈希多样性(p=0.0039;部分为操纵性检验)。因子设计未发现正向的记忆-排斥交互作用(界限约为±0.04);增益可加性分解,且记忆+排斥是唯一从未崩溃的处理组(18次运行中0次崩溃),与完整组合相差在0.025以内。在相同循环下,前沿提出者(frontier proposer)在数十个样本内达到了本地模型在数百个样本内未能达到的结果;在单次范围界定运行中,其增益在无操作符的情况下出现。搜索在旗舰问题上缩小约92%的差距后停滞,注册的家庭提示测试(family-hint test)首次解读了这一停滞:以文字命名时,参考家庭被采纳但失败;以代码形式给出时,它被优化,但我们最佳的有限网格实现仍低于无辅助情况下达到的平台。循环传递并优化了给予它的想法;没有无辅助运行产生该想法。我们发布了该框架、每个候选方案以及带日期的预注册。

英文摘要

Verified search, in which a language model proposes programs, a hard evaluator scores them, and selection keeps the best, has recently moved mathematical records; controlled ablations of the proposer-side components remain rare. We instrument a minimal FunSearch-style loop at office scale (a 30B local model on a laptop, 120-600 verified samples per run) with three operator packages: a schematic notebook the model writes and carries instead of verbatim elites, a named obstacle, and behavioural repulsion from constructions already found. On nine construction problems from a public repository, the complete 2^3 factorial with two replicates favours the primary contrast in a nominal two-stage analysis: the composition closes more of the seed-to-record gap (+0.196; nominal pooled p=0.023, stage-combination p~0.08; median per-problem effect +0.045). Repulsion raises construction-hash diversity everywhere (p=0.0039; partly a manipulation check). The factorial finds no positive memory-by-repulsion interaction (bounded to about +/-0.04); the gain decomposes additively, and memory+repulsion is the only arm that never collapses (0 of 18 runs), within 0.025 of the full composition. A frontier proposer under the identical loop reaches in tens of samples what the local model does not in hundreds; in single scoping runs its gains arrive without the operators. The search stalls after closing ~92% of the gap on the flagship problem, and the registered family-hint test gives the stall its first reading: named in words, the reference family is adopted and loses; handed as code, it is optimized, but our best finite-grid implementation remains below the plateau reached unaided. The loop transported and optimized the idea it was handed; no unaided run produced it. We release the harness, every candidate, and the dated pre-registrations.

补充信息

↑