发表机构
University of California, Santa Cruz; IBM Research(加州大学圣克鲁兹分校; IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出aLoRA适配器与上下文感知路由的模块化框架,在生成中定向消除LLM的有害输出,提升安全性且保持任务性能。
AI 中文摘要
大型语言模型(LLMs)是强大的零样本学习者,但仍容易与人类偏好不一致,经常产生有偏见、有毒或其他有害的输出。现有的对齐方法虽然有效,但成本高昂且与模型紧密耦合,限制了灵活性和可扩展性。我们提出了一种模块化校正框架,通过激活LoRA(aLoRA)适配器和上下文感知路由机制增强预训练LLMs,以消除不对齐模型响应中的危害。我们的方法使专家适配器能够在序列中激活,而不会使KV缓存失效,从而在生成过程中实现低延迟、定向校正。每个专家都经过训练,以检测和缓解特定危害,如偏见或毒性。一个学习路由器根据模型的中间输出动态选择合适的专家。我们证明,我们的系统在标准安全基准上提高了对齐性,同时保持了任务性能,为更安全、更可控的LLM部署提供了一条轻量级且高效的路径。
英文摘要
Large Language Models (LLMs) are powerful zero-shot learners but remain prone to misalignment with human preferences, often producing biased, toxic, or otherwise harmful outputs. Existing alignment methods, while effective, are costly and tightly coupled to the model, limiting flexibility and scalability. We propose a modular correction framework that augments pretrained LLMs with Activated LoRA (aLoRA) adapters and a context-aware routing mechanism to eliminate harms from misaligned model responses. Our approach enables expert adapters to activate mid-sequence without invalidating the KV cache, allowing low-latency, targeted correction during generation. Each expert is trained to detect and mitigate specific harms, such as bias or toxicity. A learned router dynamically selects appropriate experts based on the models intermediate outputs. We demonstrate that our system improves alignment on standard safety benchmarks while preserving task performance, offering a lightweight and efficient path toward safer and more controllable LLM deployments.