语义纤维与跨文法干扰:过完备表示中安全漂移的演算
Semantic Fibers and Cross-Gram Interference: A Calculus of Safety Drift in Overcomplete Representations
浏览论文内容
中文总结 AI 辅助
本文通过线性代数演算形式化跨语言安全漂移,提出暴露度量与三区域诊断,并验证了控制接口残差的高预测性能。
中文摘要 AI 辅助
部署的语言模型可能拒绝英文的有害请求,却遵从忠实翻译后的请求,这揭示了跨语言安全失败,仅凭输出行为无法可靠地表征。我们通过一个经审计的等价关系形式化这一现象,并表明对于声明的商、表示、度量、特征字典、评分头、阈值和对比模型,所产生的安全漂移具有精确的线性代数表征。具体而言,漂移是纤维内对比的跨文法泛函;其最坏可接受值是支撑函数,而边际不变性由零化子条件表征。我们引入一种内在的校准暴露度量,由杠杆对偶性 $\chi^2=1/\ell-1$ 控制,该度量将观察到的漂移分为三个诊断上不同的区域:可通过重新校准移除的读取器故障、过于病态而不可靠的精确校正,以及任何仅读取干预都无法移除的表征级碰撞。因此,相同的观察暴露可能导致根本不同的修复裁决。该框架还扩展到锥值安全头。一个解开的顺序交换恒等式为线性控制接口提供了诊断;其校准状态残差在未见状态和目标上预测出不同的三控制组合误差,中位斯皮尔曼相关系数为 0.964,而静态跨文法基线为 0.269。等。
英文摘要
A deployed language model may refuse a harmful request in English yet comply with its faithful translation, revealing a cross-lingual safety failure that cannot be characterized reliably by output behavior alone. We formalize this phenomenon through an audited equivalence relation and show that, for a declared quotient, representation, metric, feature dictionary, scoring head, threshold, and contrast model, the resulting safety drift admits an exact linear-algebraic characterization. Specifically, the drift is a cross-Gram functional of the within-fiber contrast; its worst admissible value is a support function, while margin invariance is characterized by an annihilator condition. We introduce an intrinsic calibrated exposure measure, governed by the leverage duality $χ^2=1/\ell-1$, which separates observed drift into three diagnostically distinct regimes: a reader fault removable by recalibration, an exact correction that is too ill-conditioned to be reliable, and a representation-level collision that no readout-only intervention can remove. Thus, identical observed exposure can lead to fundamentally different remediation verdicts. The framework also extends to cone-valued safety heads. An untied order-swap identity provides a diagnostic for the linear control interface; its calibration-state residual predicts a distinct three-control composition error on unseen states and targets, achieving median Spearman correlation $0.964$, compared with $0.269$ for a static cross-Gram baseline. etc.....
发表机构
- Université Paris 1(巴黎第一大学)
- Faculty of Science and Technology of Tangier(丹吉尔科学与技术学院)
- Abdelmalek Essaadi University(阿卜杜勒马利克·埃萨迪大学)
机构由 AI 辅助整理,请以论文原文为准。