为何第三轴是自由
Why the Third Axis Is Freedom
- The Australian National University(澳大利亚国立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究指出生成式训练中XM的第三轴实际是自由,通过理论与实验证明自由选择优于MDL,且在分布偏移下针对自由度的选择可提升XM表现。
AI中文摘要:
在生成式训练中,模型会生成一个输出,并因该输出与示例的差异而受到惩罚。每次比较仅对应一个输出时,生成单一常见答案的模型会优于保留更广泛输出集合的模型。探索性建模(XM)每次比较会生成K个输出,并基于最接近的输出进行更新,其声称探索是与生成表达性相关的“第三预训练轴”。本文中,我证明该第三轴实际是自由,即模型行为所隐含约束的薄弱性。此前研究表明,自由是函数的属性而非形式的属性:参数、架构、最小描述长度(MDL)和数据均可变化,而行为约束保持不变。已有正式证明,最弱的模型最有可能泛化,且在归纳实验中,自由选择的表现比MDL高出110%至500%。我证明平均XM损失取决于候选者错过可接受区域的概率,而探索会将错过概率提升至K次方;当K>1时,匹配概率随自由度上升。随后,我通过实验证明XM优化的是自由度:在正向XM实验中,更大的K会提升或饱和测得的自由度,且在所有测试的上下文依赖目标值下,自由度均随K增大而提升;我训练了XM候选池,并将验证集选择与读取未标记父上下文的自由选择器进行比较,自由选择在30次实验中赢了29次。生成表达性是自由度的模式计数代理,它丢弃了赋予自由度泛化意义的扩展结构;XM是手段,自由是目标,在分布偏移下,针对自由度的选择提升了XM的表现。
英文摘要:
In generative training, a model produces an output and is penalised for its difference from an example. With one output per comparison, a model that produces one common answer can outperform a model retaining a broader repertoire. Explorative Modeling (XM) produces $K$ outputs per comparison and updates on the closest, claiming exploration as a "third pretraining axis" associated with generative expressivity. Here I show the third axis is actually freedom, meaning the weakness of the constraint implied by a model's behaviour. Previous work showed freedom is a property of function rather than form. Parameters, architecture, minimum-description-length (MDL), and data can vary while the behavioural constraint remains unchanged. It was formally proved that weakest models are likeliest to generalise, and freedom selection beat MDL by 110-500\% in induction experiments. I prove average XM loss depends on the chance a candidate misses an acceptable region, with exploration raising miss probability to power $K$. For $K>1$, match probability rises with freedom. I then demonstrate empirically that XM optimises for freedom. In a Forward XM experiment, larger $K$ increased or saturated measured freedom, and increased freedom at every tested value under context-dependent targets. I trained XM candidate pools and compared validation selection with a freedom selector that read unlabelled parent contexts. Freedom won in 29 of 30 cases. Generative expressivity is a mode-count proxy for freedom, that discards the extension structure that gives freedom its generalisation significance. XM is a means, freedom an end, and selecting for freedom improved XM under distribution shift.