arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于LLMS中定向危害缓解的高效模块化框架

An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS

Roberto Campbell, Momin Abbas, Muneeza Azmat, Michal Ulewicz, Raya Horesh, Kristjan Greenewald, Rogério Abreu de Paula, Nathalie Baracaldo

arXiv 2609.13624首次发表:更新:

发表机构

University of California, Santa Cruz; IBM Research(加州大学圣克鲁兹分校; IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出aLoRA适配器与上下文感知路由的模块化框架,在生成中定向消除LLM的有害输出,提升安全性且保持任务性能。

AI 中文摘要

大型语言模型(LLMs)是强大的零样本学习者,但仍容易与人类偏好不一致,经常产生有偏见、有毒或其他有害的输出。现有的对齐方法虽然有效,但成本高昂且与模型紧密耦合,限制了灵活性和可扩展性。我们提出了一种模块化校正框架,通过激活LoRA(aLoRA)适配器和上下文感知路由机制增强预训练LLMs,以消除不对齐模型响应中的危害。我们的方法使专家适配器能够在序列中激活,而不会使KV缓存失效,从而在生成过程中实现低延迟、定向校正。每个专家都经过训练,以检测和缓解特定危害,如偏见或毒性。一个学习路由器根据模型的中间输出动态选择合适的专家。我们证明,我们的系统在标准安全基准上提高了对齐性,同时保持了任务性能,为更安全、更可控的LLM部署提供了一条轻量级且高效的路径。

英文摘要

Large Language Models (LLMs) are powerful zero-shot learners but remain prone to misalignment with human preferences, often producing biased, toxic, or otherwise harmful outputs. Existing alignment methods, while effective, are costly and tightly coupled to the model, limiting flexibility and scalability. We propose a modular correction framework that augments pretrained LLMs with Activated LoRA (aLoRA) adapters and a context-aware routing mechanism to eliminate harms from misaligned model responses. Our approach enables expert adapters to activate mid-sequence without invalidating the KV cache, allowing low-latency, targeted correction during generation. Each expert is trained to detect and mitigate specific harms, such as bias or toxicity. A learned router dynamically selects appropriate experts based on the models intermediate outputs. We demonstrate that our system improves alignment on standard safety benchmarks while preserving task performance, offering a lightweight and efficient path toward safer and more controllable LLM deployments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑