arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

细粒度危害信号在LLM安全中的作用

The Role of Fine-grained Harm Signals in LLM Safety

Soyeon Park, Seogyeong Jeong, Sunwoo Kim, Alice Oh

arXiv 2609.19366首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过分离LLM中类别特异性危害表征,发现细粒度类别残差在编码危害性和引发拒绝上因类别而异,且能增强下游通用危害对齐,强调除通用表征外需考虑细粒度信号以全面理解LLM安全。

AI 中文摘要

先前的研究表明,大型语言模型中的内部危害性表征在不同风险类别间存在差异,同时共享一个通用的危害表征组件。这引发了一个问题:在LLM安全中,类别特定组件在通用危害表征之外扮演什么角色?为了回答这个问题,我们通过从每个类别危害表征中移除共享的通用危害性表征来隔离类别特定组件,从而在每个层上产生一个与通用危害性正交的类别残差。通过在3个指令微调的LLM中对11个风险类别使用类别残差进行激活引导,我们发现类别残差是否编码危害性因类别而异,且这种类别模式在不同模型间相似。类别残差是否引发拒绝也因类别而异,但这种类别模式更依赖于模型。我们还发现,类别残差增强了LLM下游与共享通用危害性表征的内部对齐。综合这些发现表明,为了全面理解LLM安全,除了共享的通用危害性表征外,还应考虑更细粒度的类别残差。更广泛地说,我们的发现表明,即使在一个层上与概念正交的方向也能对该概念的下游放大做出贡献。

英文摘要

Prior work has shown that internal harmfulness representations in large language models vary across risk categories, while sharing a common general harm representation component. This raises a question about the role of the category-specific component beyond general harm representation in LLM safety. To answer this question, we isolate the category-specific component by removing shared general harmfulness representation from each categorical harmfulness representation, yielding a category residual that is orthogonal to general harmfulness at every layer. Using activation steering with category residuals across 11 risk categories in 3 instruction-tuned LLMs, we find that whether category residuals encode harmfulness varies across categories, and that this category-wise pattern is similar across models. Whether category residuals induce refusal also varies across categories, but this category-wise pattern is more model-dependent. We also find that category residuals increase LLMs' downstream internal alignment with shared general harmfulness representation. Together, these findings demonstrate that more fine-grained category residuals should also be considered beyond shared general harmfulness representation to fully understand LLM safety. More broadly, our findings show that even a direction orthogonal to a concept at one layer can contribute to the concept's downstream amplification.

Comments9 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑