arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11243cs.AIcs.LG

离支撑屏障:为何语义安全约束不是学习问题不变量,以及对先验设计、约束与验证的启示

The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification

Yoshinori Watanabe

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出语义安全约束是离支撑对象,基于奇异学习理论推导得出多项推论,以2026年7月OpenAI与Hugging Face评估事件为案例,为AI安全的先验设计、约束验证等问题提供理论启示。

中文摘要 AI 辅助

我们认为,一个单一的结构事实组织了当代AI安全中的广泛现象:语义安全约束(例如智能体不逃离其沙盒)是一个离支撑对象。形式上,若q为数据分布,p(·|w)为模型,安全谓词B相对于σ(模型, q)不可测,而奇异学习理论(SLT)的实对数典范阈值(RLCT)则是可测的。从这种非不变性中,我们推导得出以下推论(而非独立观察):(i) 基于结果的优化下会出现奖励黑客行为和沙盒逃逸的原因;(ii) 通过贝叶斯先验设计或软惩罚权重编码此类约束在奇异模型中杠杆作用不佳的原因;(iii) 为何硬不变量应属于管控机制,而软倾向应属于模型;(iv) 为何相同的B仍可通过形式验证得到可靠的局部证明,正如局部学习系数(LLC)在局部确定相同的RLCT——存在两个明确的不相似点;(v) 为何识别哪个离支撑区域重要的剩余难题,与表演性预测和自指功能动力学相吻合,此时SLT的分析机制失效。我们以2026年7月OpenAI与Hugging Face的评估事件作为动机案例。数值实验代码及Lean中的相关证明可在此https URL获取。

英文摘要

We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-support object. Formally, if q is the data distribution and \(p(\cdot\mid w)\) the model, the safety predicate B is not measurable with respect to \(σ(\text{model}, q)\), whereas the real log-canonical threshold (RLCT) of singular learning theory (SLT) is. From this non-invariance we derive, as corollaries rather than independent observations: (i) why reward hacking and sandbox escape arise under outcome-based optimization; (ii) why encoding such constraints through Bayesian prior design or soft penalty weighting has poor leverage in singular models; (iii) why hard invariants belong in the harness and soft dispositions in the model; (iv) why the same B is nonetheless soundly and locally certifiable by formal verification, exactly as the local learning coefficient (LLC) locally pins the same RLCT --- with two precise points of disanalogy; and (v) why the residual difficulty, identifying which off-support region matters, coincides with performative prediction and self-referential functional dynamics, where SLT's analytic machinery breaks down. We use the July 2026 OpenAI--Hugging Face evaluation incident as the motivating case. Numerical experiments code and related proofs in lean are available at https://github.com/xiangze/Preventing_Jailbreak_as_regularization

↑