arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于模型选择的贝叶斯风洞

Bayesian Wind Tunnels for Model Selection

Siddhartha R Dalal, Vishal Misra, Abhay Parekh

arXiv 2607.19379首次发表:更新:

发表机构

Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究变压器能否进行贝叶斯模型选择,引入模型选择贝叶斯风洞。通过无不动点对合,变压器在特定环境下实现与贝叶斯最优值的熵一致性,拓展到非嵌套比较,还确定了感知访问条件及相关因素,展示了模型选择能力并揭示了语言模型的校准问题。

AI 中文摘要

先前的工作表明,变压器可以在固定假设类中执行精确的贝叶斯滤波。那么它们能否执行贝叶斯模型选择——从数据中识别正确的假设类呢?我们引入了模型选择贝叶斯风洞:在封闭形式下可获得假设类的真实后验的受控环境。使用无不动点对合(其定义属性\(f(f(x)) = x\)纯粹是关系性的),一个具有280万个参数的变压器与贝叶斯最优值实现了0.01比特的熵一致性(3个种子),无论是整数令牌还是含义每集都变化的不透明符号。这扩展到非嵌套比较:对合与3 - 循环(其中一个类不是另一个类的子集)实现了低于0.001的类后验平均绝对误差,展示了超越简单性/子集偏差的真正模型选择。然后我们确定了一个敏锐的感知访问条件:当判别统计需要算术运算——模加法(旋转)或乘法(\(f(x)=cx \bmod p\))时,模型选择对于整数令牌成功,但对于不透明符号完全失败,并且在112倍缩放(280万到3.16亿参数)下这种界限仍然存在。平稳性控制证实了操作因素:具有固定重新标记的不透明令牌成功(0.009比特平均绝对误差),表明稳定的语义而非整数标识能够实现电路编译。头部子任务诊断将失败定位到头部反转与算术运算的组合而非头部解析本身。在相同任务上探测前沿语言模型显示出定性的贝叶斯行为,但存在较大的校准差距(约55倍),通过有损探测测量,因此是方向性的而非精确的。

英文摘要

Prior work has shown that transformers can perform exact Bayesian filtering within a fixed hypothesis class. Can they also perform Bayesian model selection -- identifying the correct hypothesis class from data? We introduce model-selection Bayesian wind tunnels: controlled environments where ground-truth posteriors over hypothesis classes are available in closed form. Using fixed-point-free involutions -- whose defining property f(f(x))=x is purely relational -- a 2.8M-parameter transformer achieves 0.01-bit entropy agreement with the Bayesian optimum (3 seeds), with both integer tokens and opaque symbols whose meanings change every episode. This extends to non-nested comparisons: involutions vs. 3-cycles (where neither class is a subset of the other) achieve class-posterior MAE under 0.001, demonstrating genuine model selection beyond simplicity/subset bias. We then identify a sharp perceptual access condition: when the discriminative statistic requires arithmetic -- modular addition (rotations) or multiplication (f(x)=cx mod p) -- model selection succeeds with integer tokens but fails completely with opaque symbols, and this boundary persists under 112x scaling (2.8M to 316M parameters). A stationarity control confirms the operative factor: opaque tokens with a fixed relabeling succeed (0.009-bit MAE), showing that stable semantics, not integer identity, enable circuit compilation. Header subtask diagnostics localize the failure to the composition of header inversion with arithmetic rather than header parsing itself. Probing frontier LLMs on the same tasks shows qualitative Bayesian behavior but a large calibration gap (~55x), measured through lossy probes and therefore directional rather than exact.

Comments27 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑