arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

走向对齐缩放定律:一个框架与首次预注册测量

Toward Alignment Scaling Laws: A Framework and First Preregistered Measurements

Jeremy Canale

arXiv 2610.08540首次发表:更新:

AI 中文总结

本文提出对齐缩放定律框架,将对齐负担建模为能力代理的幂律,证明最大指数决定长期状态,并通过预注册实验发现某些风险缩放有帮助,后门问题仍存疑。

AI 中文摘要

对齐在模型规模增大时是变得更容易还是更困难,通常是从孤立的研究结果中论证的,仿佛对齐是一个单一属性。我们将其视为一族可测量的缩放关系:对于每个风险类别r,将安全性目标保持在固定水平所需的对齐负担建模为B_r(N)=a_rN^alpha_r,其中N是能力代理指标;相对于与N成比例的预算,如果alpha_r<1,则缩放有助于对齐;如果alpha_r~1,则保持同步;如果alpha_r>1,则积累对齐债务。我们给出了负担的三种操作性定义,并区分了观察到的、经审计的和真实的对齐。一个玩具模型(其中修正消耗能力余量)使后果变得明确。我们证明了被修正风险中的最大指数(而非平均值)决定了长期状态;当指数高于1时,任何将余量保持在底线以上的策略都必须超指数增长;对于正混合幂律的负担,小模型上的拟合会低估大规模指数;而一个能发现隐藏失败且无误报的审计永远不会低估真实对齐。我们提出一个可预注册的协议,并对其简化版本应用了两次。对Pythia分类器公开对抗训练数据的预注册再分析发现,将攻击成功率降至10%以下所需的计算量随N^0.60增长。在Qwen2.5 0.5B-72B上的预注册试点发现,真实性指数为-0.05,陈述倾向指数为0.48(在其简化规则下两者均显示缩放有帮助,尽管在最大规模处局部斜率接近1;在Qwen3 0.6B-14B上得到复现),而谄媚(0.89,或在72B处添加两个种子后为0.83)和植入的后门则无法确定:后门在触发器已知时被快速移除,但在五个规模中的四个上,盲安全性训练后仍存活。我们发布了四个浏览器游戏来演示这些定律(此http URL)。我们不声称当前前沿模型处于哪种状态。

英文摘要

Whether alignment gets easier or harder as models grow is often argued from isolated findings, as if alignment were one property. We treat it as a family of measurable scaling relations: for each risk category r, the alignment burden needed to hold a fixed safety target is modeled as B_r(N)=a_rN^alpha_r, with N a capability proxy; against a budget proportional to N, scaling helps if alpha_r<1, keeps pace if alpha_r~1, and accumulates alignment debt if alpha_r>1. We give three operationalizations of burden and distinguish observed, audited and true alignment. A toy model, in which corrections consume capability headroom, makes the consequences explicit. We prove that the largest exponent among corrected risks, not an average, sets the long-run regime; that above 1 any policy holding headroom above a floor must grow super-exponentially; that, for burdens that are positive mixtures of power laws, fits on small models underestimate large-scale exponents; and that an audit that uncovers hidden failures without false positives never underestimates true alignment. We propose a pre-registrable protocol and apply reduced versions of it twice. A preregistered reanalysis of public adversarial-training data for Pythia classifiers finds that the compute needed to bring attack success under 10% grows as N^0.60. A preregistered pilot on Qwen2.5 0.5B-72B finds exponents of -0.05 for truthfulness and 0.48 for stated dispositions (both scaling helps under its reduced rule, though local slopes approach 1 at the top; replicated on Qwen3 0.6B-14B), while sycophancy (0.89, or 0.83 with two seeds added at 72B) and a planted backdoor are undetermined: the backdoor is removed quickly when its trigger is known but survives blind safety training at four of five sizes. We release four browser games that play these laws (www.aisafety.fun). We make no claim about which regime holds for current frontier models.

Comments34 pages, 24 figures, 8 tables. Games: https://www.aisafety.fun. Preregistrations: https://osf.io/wda8q, https://osf.io/q2j3y, https://osf.io/8kreb

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑