发表机构
Lanzhou University; National University of Singapore; Beijing Institute of Technology(兰州大学; 新加坡国立大学; 北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DNAlign,一种结合控制理论优化与零空间投影的轻量级LLM安全对齐框架,通过可控扰动和投影模块在减少有害输出的同时保持知识、流畅性与事实性,优于现有基线。
AI 中文摘要
确保大型语言模型(LLMs)的安全可靠部署仍然是一个基本挑战。现有的安全对齐方法要么带来高昂的计算成本,要么无意中破坏模型的核心知识,导致在良性任务上流畅性和事实准确性下降。这揭示了安全性与实用性之间持续存在的权衡。我们提出了DNAlign,一个轻量级对齐框架,将控制理论优化与零空间投影相结合。通过将LLM视为动态系统,所提出的框架引入可控扰动以引导生成朝向安全行为。关键组件是投影模块,该模块将这些扰动限制在从中性隐藏状态导出的有害相关子空间中,从而保留一般知识和响应质量。基于人类偏好数据训练的价值函数自适应地优化控制信号,以符合人类安全偏好。在多个LLM骨干上的广泛评估表明,我们的框架持续减少有害输出,同时保持流畅性、连贯性和事实效用。与先前的对齐基线相比,它在不牺牲生成多样性的情况下实现了优越的整体性能。这些结果表明,所提出的框架为安全LLM对齐提供了一种有效且可实际部署的解决方案。代码可在以下网址获取:此https URL。
英文摘要
Ensuring the safe and reliable deployment of large language models (LLMs) remains a fundamental challenge. Existing safety alignment approaches either incur high computational cost or unintentionally disrupt the model's core knowledge, leading to degraded fluency and factual accuracy on benign tasks. This reveals a persistent trade-off between safety and utility. We propose DNAlign, a lightweight alignment framework that integrates control-theoretic optimization with null-space projection. By treating the LLM as a dynamic system, the proposed framework introduces controllable perturbations to steer generation toward safe behavior. A key component is the projection module, which restricts these perturbations to the harmful-related subspace derived from neutral hidden states, thereby preserving general knowledge and response quality. A value function trained on human preference data adaptively optimizes the control signals to align with human safety preferences. Extensive evaluations across multiple LLM backbones demonstrate that our framework consistently reduces harmful outputs while maintaining fluency, coherence, and factual utility. It achieves superior overall performance compared to prior alignment baselines without sacrificing generation diversity. These results indicate that the proposed framework provides an effective and practically deployable solution for safe LLM alignment. Code is available at https://anonymous.4open.science/r/DNAlign.
CommentsRegular Paper; 13 pages, 8 figures, and 1 table