AI 中文总结
该研究以在线装箱问题为对象,发现经价值筛选候选训练及LoRA整合可使生成模型的最优输出收敛于经典启发式水平,均值可稳定复制,为开放域生成的推理时干预与权重整合提供了实证依据。
AI 中文摘要
从生成循环中学习自身已验证的成功能获得什么?在在线装箱问题的生成、验证、选择与LoRA整合循环中,对经价值筛选的候选进行训练,会使模型在保留变体上的输出向价值指标偏移(超额值为-1.7,p=0.008;与随机整合对照组相比为-3.1,p=0.004),而观测到的最优候选收敛至经典启发式算法水平,且不再进一步提升。验证组重复整个流程三次,使用新的随机种子及仅查询一次的未接触保留集:三个谱系的均值几乎相同(-2.0、-1.8、-1.9),在所有7个可评估变体中,按保留变体聚合后均支持整合(p=0.008)。观测到的最优候选精确达到经典启发式算法水平(三个谱系中均为0.021028,吸引组与随机对照组一致),且从未超越该水平。匹配的仅SFT对照组显示,是监督锚点而非对坏候选的排斥导致了集中效应(96%的候选精确落在经典启发式算法水平)。存在双向尾部效应:整合降低了优于经典算法的候选比例(从10%降至3.9%),但更大的产出规模使其绝对数量更多(少数事件中为5对1)。作为动机,我们报告了导向此结果的推理时台账:模型生成的概要图有助于评判文档整合,无内容有助于开发;写入流中的验证器被模仿,每本笔记本产生16.4条伪造判定行。有效候选的平均质量可被获取并复制;观测到的最优候选达到经典水平,且目前从未超越。
英文摘要
What does a generation loop gain from learning on its own verified successes? In cycles of generate, verify, select and LoRA-consolidate on online bin packing, training on value-filtered candidates shifts what the model writes on held-out variants toward value (-1.7 points of excess, p=0.008; -3.1 against a random-consolidation control, p=0.004) while the best observed candidate converges to the classic heuristic's level and no further. A confirmation battery replicates the whole procedure three times, with fresh seeds and a never-consulted held-out set read exactly once: the mean was nearly identical in all three lineages (-2.0, -1.8, -1.9), and after aggregating within held-out variant all seven evaluable variants favored consolidation (p=0.008). The best observed candidate moved to the classic heuristic's level, exactly (0.021028 in all three lineages, for attract and for the random control alike), and never beyond it. A matched SFT-only control shows the supervised anchor, not repulsion from bad candidates, does the concentrating (96% of candidates land exactly at the classic heuristic's level). The tails cut both ways: consolidation lowers the per-candidate rate of better-than-classic candidates (10% to 3.9%) while its larger production yields more such candidates absolutely (5 against 1, on few events). As motivation we report the inference-time ledger that led here: a model-written schematic recap buys judged document integration and nothing buys development; a verifier written into the stream is imitated, 16.4 fabricated verdict lines per notebook. Mean quality among valid candidates can be bought and replicated; the observed best goes to the classic and, so far, never beyond it.
CommentsCode and run data: https://github.com/RobertoOno/interrupting-the-loop