arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

规则还是特性?AI安全设计的缩放定律

Rules or Character? Scaling Laws for AI Safety Design

Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu, Ryuji Hamamoto

arXiv 2608.13345首次发表:更新:

AI 中文总结

该研究构建模型分析AI安全的特性塑造与规则执行的资源分配,发现部署规模对最优分配影响较小,特性脆弱率才是主导因素,二者的最优值在大规模部署下会收敛。

AI 中文摘要

人工智能(AI)安全系统将特性塑造(例如人类反馈强化学习[RLHF]、宪法AI)与规则执行(例如输出过滤器、安全分类器)相结合,前者在训练时修改行为分布,后者在推理时阻止有害输出,但关于随着部署规模增加,二者的最优平衡应如何变化,目前几乎没有正式分析。我们引入了一个程式化的比较静态模型,将安全设计参数化为在[0,1]之间的资源分配α,分配给上述两种方法,该模型纳入了与规模相关的过滤器退化、共模故障以及特性脆弱性——即经塑造的行为在新条件下退化或崩溃的风险。在乘法帕累托损伤模型下,我们推导了闭式期望损伤,并通过蒙特卡洛模拟补充了尾部风险(条件风险价值CVaR)分析。在三种场景(乐观、中等、悲观)下,最优α*为内部值或仅规则边界,且随着部署规模T增大,会微弱地向特性塑造方向偏移,偏移量从可忽略(Δα*=+0.01)到显著(Δα*=+0.21),具体取决于场景。主导参数是基线特性脆弱率p^(0)_frag,其在自身取值范围内可使α*偏移0.50——远超过尾部严重程度、过滤器质量或共模故障概率的影响。在大规模T下,CVaR和期望损伤的最优值会收敛。这些结果表明,安全架构决策更多取决于分布转移下特性塑造的可靠性,而非部署规模本身。

英文摘要

Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distributions at training time, with rule enforcement (e.g., output filters, safety classifiers), which blocks harmful outputs at inference time, yet little formal analysis exists on how their optimal balance should change as deployment scales increase. We introduce a stylized comparative-statics model that parameterizes safety design as a resource allocation alpha in [0,1] between these two approaches, incorporating scale-dependent filter degradation, common-mode failures, and character fragility -- the risk that shaped behavior degrades or collapses under novel conditions. Under a multiplicative Pareto damage model, we derive closed-form expected harm and supplement it with tail-risk (CVaR) analysis via Monte Carlo simulation. Across three scenarios (optimistic, moderate, pessimistic), the optimal alpha* is interior or at the rules-only boundary and shifts weakly toward character shaping as deployment scale T grows, from negligible (Delta alpha* = +0.01) to pronounced (Delta alpha* = +0.21) depending on scenario. The dominant parameter is the baseline character fragility rate p^(0)_frag, which shifts alpha* by 0.50 across its range -- far exceeding the effect of tail severity, filter quality, or common-mode failure probability. CVaR and expected-harm optima converge at large T. These results suggest that safety architecture decisions depend less on deployment scale per se than on the reliability of character shaping under distributional shift.

CommentsAccepted at AIES 2026 (9th AAAI/ACM Conference on AI, Ethics, and Society). 10 pages, 6 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑