可靠性覆盖中的Oracle差距:采样噪声还是策略特化?
Oracle Gaps in Reliability Coverage: Sampling Noise or Policy Specialization?
浏览论文内容
中文总结 AI 辅助
本研究通过可靠性覆盖检验oracle差距的来源,发现采样噪声而非策略特化主导,且路由器增益仅在特化明显时出现。
中文摘要 AI 辅助
从同一基础模型训练出的策略可能看起来在解决不同的问题。因此,为每个问题选择最佳策略的oracle可能看起来比任何单一策略都强大得多。选择最大的估计成功率也会选择有利的采样误差。我们通过可靠性覆盖(即成功概率达到选定阈值的问题比例)来研究这一效应。我们的第一个测试在每个问题内跨策略重新分配存储的正确性结果。第二个测试还保留每个策略的总成功数,在指定统计模型下考虑整体质量差异。对于七亿参数视觉语言模型的五个训练种子,在阈值0.10下,重新分配重现了估计的0.113 oracle差距中的0.096。两个测试在该阈值下均未发现显著证据。小的优势仍未解决。混合、路由器、投票和权重平均相对于相应的单一策略基线没有显示出可检测的改进。在轻量后训练下,在不同数据集上训练的策略几乎没有显示出可检测的特化,且路由器没有增益。更强的配方确实创造了特化,两个测试都检测到了它,并且路由器增益仅出现在专家分离的高阈值处。在具有可预测特化的对照中,路由器恢复了约一半的oracle差距。覆盖界限解释了为什么即使是真正的oracle优势也不一定能带来部署增益。这些测试在投资路由之前评估了存储响应中的表观特化。代码:此https URL
英文摘要
Policies trained from the same base model can appear to solve different problems. An oracle that chooses the best policy for each problem may therefore appear much stronger than any single policy. Selecting the largest estimated success rate also selects favorable sampling errors. We study this effect through reliability coverage, the fraction of problems whose success probability reaches a chosen threshold. Our first test redistributes stored correctness outcomes across policies within each problem. A second also preserves each policy's total successes, accounting for overall quality differences under a specified statistical model. For five training seeds of a seven-billion-parameter vision-language model, redistribution reproduces 0.096 of an estimated 0.113 oracle gap at threshold 0.10. Neither test finds significant evidence at this threshold. Small advantages remain unresolved. Mixtures, routers, voting, and weight averaging show no detectable improvement over their corresponding single-policy baselines. Training policies on different datasets shows little detectable specialization under light post-training, and no router gain. A stronger recipe does create it, both tests detect it, and a router gain appears only at the high thresholds where the specialists separate. In a control with predictable specialization, a router recovers about half the oracle gap. Coverage bounds explain why even a genuine oracle advantage need not yield a deployment gain. The tests assess apparent specialization from stored responses before investment in routing. Code: https://github.com/KurbanIntelligenceLab/oracle-gaps