arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于知识与安全的拒绝行为的统一机制分析

A Unified Mechanistic Analysis of Knowledge- and Safety-Based Refusals

Yuri Son, Seunghee Kim, Hyuhng Joon Kim, Taeuk Kim

arXiv 2609.00760首次发表:更新:

发表机构

Hanyang University; Samsung Electronics(汉阳大学; 三星电子)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对213组对比四元组数据集,分析发现大型语言模型的基于知识与安全的拒绝行为存在共享机制但不对称,上层存在类型特异性专业化,将拒绝过程描述为“先提交后指定”。

AI 中文摘要

大型语言模型(LLMs)接受的训练越来越多,要求它们拒绝超出自身知识范围的查询(基于知识的拒绝,KR)或违反安全政策的查询(基于安全的拒绝,SR)。尽管KR和SR会产生表面上相似的响应,但它们在很大程度上是被单独研究的,因此仍不清楚二者是否存在潜在的共同机制。我们通过对包含213组对比四元组的新数据集开展系统研究,来填补这一空白,该数据集可同时探究两种拒绝类型。我们发现KR和SR受重叠但可区分的机制支配。二者共享一个拒绝方向,但重叠程度是不对称的:SR信号向KR的转移比反向转移更强。类型特异性专业化主要出现在上层,其中KR与不确定性及知识相关的表示对齐,SR则与安全及政策相关的表示对齐。因此,我们将拒绝描述为一个“先提交后指定”的过程:共享的初始机制先确定拒绝,随后后续层的类型特异性特征会指定拒绝的依据是认知性的还是规范性的。

英文摘要

Large language models (LLMs) are increasingly trained to decline queries that fall outside their knowledge (knowledge-based refusal, KR) or violate safety policies (safety-based refusal, SR). Although KR and SR result in superficially similar responses, they have largely been studied in isolation, leaving open whether they share an underlying mechanism. We address this gap with a systematic study on a new dataset of 213 contrastive quadruples that jointly probe both refusal types. We find that KR and SR are governed by overlapping yet distinguishable mechanisms. Both share a refusal direction, yet the overlap is asymmetric: SR signals transfer more strongly to KR than the reverse. Type-specific specialization emerges mainly in upper layers, with KR aligning with uncertainty- and knowledge-related representations and SR with safety- and policy-related ones. We thus characterize refusal as a commit-then-specify process: a shared initial mechanism commits to refusing, then type-specific features in later layers specify whether the grounds are epistemic or normative.

CommentsAccepted to EMNLP 2026 (Main Conference)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑