arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21278cs.AI

CLEAR:用于保持效用的大语言模型安全对齐的连续潜在适配器路由

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

  • Stanford University(斯坦福大学)
  • University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Siebel School of Computing and Data Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校西贝尔计算与数据科学学院)
  • Computer Science, Stanford University(斯坦福大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

Chengxiao Wang, Enyi Jiang, Xiaojing Liao, Sanmi Koyejo

AI总结:

本研究提出CLEAR框架,通过轻量型隐藏状态门控控制安全低秩适配器的激活强度,在降低LLM有害输出的同时,减少了全局安全调优带来的效用损失,实现了安全与效用的更好平衡。

AI中文摘要:

提升大语言模型(LLM)的安全性往往会以牺牲效用为代价,因为全局应用的安全调优可能会影响模型对有害和良性输入的响应。我们提出了连续潜在适配器路由(CLEAR),这是一种条件安全适配框架,它使用轻量型隐藏状态门控来连续控制安全低秩适配器的激活强度。CLEAR旨在减少有害输出,同时避免对冻结的主干模型进行不必要的修改,这些修改可能会降低模型在良性提示上的性能。在广泛使用的安全和效用基准上进行的实验表明,CLEAR提高了HarmBench的鲁棒性,同时降低了全局应用安全调优(如SFT或标准低秩适配(LoRA))所观察到的效用下降。在Llama-3-8B-Instruct上,CLEAR将HarmBench的攻击成功率(ASR)从32.3%降至0.5%,同时保留了基础模型的大部分效用,并且比全局应用的SFT或LoRA实现了高达7.1个百分点的GSM8K准确率提升。这些结果表明,CLEAR是改善LLM对齐中安全-效用权衡的一种有前景的机制。

英文摘要:

Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safety tuning may affect model responses to both harmful and benign inputs. We propose \textbf{C}ontinuous \textbf{L}at\textbf{E}nt \textbf{A}dapter \textbf{R}outing (CLEAR), a conditional safety adaptation framework that uses a lightweight hidden-state gate to continuously control the activation strength of a safety low-rank adapter. CLEAR aims to reduce harmful completions while avoiding unnecessary changes to the frozen backbone that could degrade performance on benign prompts. Experiments on widely used safety and utility benchmarks show that CLEAR improves robustness on HarmBench while reducing the utility degradation observed with globally applied safety tuning such as SFT or standard low-rank adaptation (LoRA). On Llama-3-8B-Instruct, CLEAR reduces HarmBench ASR from 32.3\% to 0.5\%, while retaining most of the base model's utility and achieving up to 7.1 percentage points higher GSM8K accuracy than globally applied SFT or LoRA. These results suggest that CLEAR is a promising mechanism for improving the safety--utility trade-off in LLM alignment.

↑