Laya 作为类型化概率评估器:一项独立复现及校准与选择性升级的预注册研究
Laya as a Typed Probabilistic Assessor: An Independent Reproduction and a Preregistered Study of Calibration and Selective Escalation
浏览论文内容
中文总结 AI 辅助
本研究独立复现并预注册评估 Laya 类型化概率评估器,发现其整体置信度不足,通过温度缩放有效校准,而等渗回归失败,且选择性升级未达错误目标,无零样本迁移。
中文摘要 AI 辅助
已发布的 Laya Typed-Decisions 检查点是一个 421M 参数的 ModernBERT-large 评估器,用于对工作流状态上的类型化 choice/noul/score 问题进行回答,该检查点整体上置信度不足。有符号的置信度-准确率差距为 $-0.214$,每个被占用的可靠性分箱的准确率都超过其置信度,且这种符号一致性将所有分箱 ECE 变体折叠为相同的值 $0.214$。模型卡片将风险描述为过度自信;而实测方向恰恰相反,该方向决定了置信度门控级联的失效方式。单一不相交拟合的温度($T=0.469$,锐化)消除了大部分校准误差(留出 ECE 从 $0.204$ 降至 $0.037$),并优于已发布的按选项计数表。冻结的选择规则反而选择了等渗回归,该回归在两个轨道上均过拟合且未通过其留出 NLL 对比,因此假设 H2 不被支持。在完整的官方测试划分上重新运行已发布的检查点,复现了模型卡片的标题准确率($0.767$ 对比 $0.766$)。回顾性 E1 复现先于分析冻结;E2-E8 为前瞻性预注册,22 项已执行的确认性检验中有 20 项在 Benjamini-Hochberg FDR 下以 $q=0.05$ 被拒绝(两项被降级)。冻结的门控优于随机升级,但在两个轨道上均未达到其 10% 的接受集错误目标,一项探索性分布外探测发现无零样本迁移(准确率 $0.617$),且每个分数衡量的是与一个合成教师的一致性,该教师的自我一致性上限($0.735$)被专家模型超越。逐决策预测、运行清单和冻结的预注册文件均包含在辅助文件中。作者与模型发布方、数据集发布方或 TypeSafe 均无关联。
英文摘要
The shipped Laya Typed-Decisions checkpoint, a 421M-parameter ModernBERT-large assessor that answers typed choice/noul/score questions over workflow state, is uniformly under-confident. The signed confidence-accuracy gap is $-0.214$, every occupied reliability bin's accuracy exceeds its confidence, and that sign uniformity collapses every binned ECE variant to the same value, $0.214$. The card frames the risk as over-confidence; the measured direction is the opposite, and the direction decides which way a confidence-gated cascade fails. A single disjointly fitted temperature ($T=0.469$, sharpening) removes most of the miscalibration (held-out ECE $0.204$ to $0.037$) and outperforms the shipped per-option-count table. The frozen selection rule instead chose isotonic regression, which overfit and failed its held-out NLL contrast on both tracks, so hypothesis H2 is not supported. Re-running the released checkpoint on its full official test split reproduces the card's headline accuracy ($0.767$ vs. $0.766$). The retrospective E1 reproduction preceded the analysis freeze; E2-E8 were prospectively preregistered, and 20 of 22 executed confirmatory tests reject under Benjamini-Hochberg FDR at $q=0.05$ (two descoped). The frozen gate beats random escalation but misses its 10% accepted-set error target on both tracks, an exploratory out-of-distribution probe finds no zero-shot transfer (accuracy $0.617$), and every score measures agreement with a synthetic teacher whose self-agreement ceiling ($0.735$) the specialist exceeds. Per-decision predictions, run manifests, and the frozen preregistration are in the ancillary files. The author has no affiliation with the model's publisher, the dataset's publisher, or TypeSafe.