arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体在真实环境中的安全上下文切换:通过正交适配缓解子空间干扰

Safe Context Switching for Agents in the Wild: Mitigating Subspace Interference via Orthogonal Adaptation

Akash Das, Ishan Roy

arXiv 2610.05219首次发表:更新:

发表机构

Fidelity Investments(富达投资)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出AURA谱正则化框架,通过强制推理与安全表征的谱独立性并约束推理更新到对齐流形正交补空间,缓解顺序子空间干扰,在恢复23.0%损失性能的同时保持安全状态余弦保真度大于0.98。

AI 中文摘要

大多数大型语言模型在两项顺序任务(如逻辑推理和安全对齐)之间表现出根本性的张力。复杂思维链(CoT)演绎所需的高方差内部状态会与编码安全约束的潜在表征产生几何干扰。我们将这一现象识别为顺序子空间干扰,表明在多步数学和代码生成等逻辑任务上进行标准微调,会在对齐基准上导致23.3%的干扰惩罚,显著削弱模型的安全先验。这种推理漂移未被当前的适配方法充分捕获,因为逻辑任务的梯度很少与安全目标正交。为解决此问题,我们提出AURA(自适应唯一残差分配),一种谱正则化框架,强制推理与安全之间的谱独立性。通过显式估计对齐流形的零空间,并将推理更新约束在其正交补空间上,AURA使模型能够在提升逻辑推理能力的同时不损害安全性。实验上,AURA恢复了23.0%的损失性能,同时保持对安全状态大于0.98的余弦保真度,证明推理与对齐可以通过几何正则化有效解耦。

英文摘要

Most Large Language Models exhibit a fundamental tension between two sequential tasks, such as logical reasoning and safety alignment. The high-variance internal states required for sophisticated Chain-of-Thought (CoT) deduction can geometrically interfere with latent representations encoding safety constraints. We identify this phenomenon as Sequential Subspace Interference, showing that standard fine-tuning on logical tasks such as multi-step mathematics and code generation can result in a 23.3% interference penalty on alignment benchmarks, substantially weakening the model's safety priors. This Reasoning Drift is not adequately captured by current adaptation methods because gradients for logical tasks are rarely orthogonal to safety objectives. To address this issue, we propose AURA (Adaptive Unique Residual Allocation), a spectral regularization framework that enforces Spectral Independence between reasoning and safety. By explicitly estimating the null space of the alignment manifold and constraining reasoning updates to its orthogonal complement, AURA enables models to improve logical reasoning without compromising safety. Empirically, AURA recovers 23.0% of the lost performance while preserving greater than 0.98 cosine fidelity to the safe state, demonstrating that reasoning and alignment can be effectively decoupled through geometric regularization.

Journal refICLR 2026 the 2nd Workshop on World Models: Understanding, Modelling and Scaling

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑