arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RISA:大语言模型拒绝校准的响应检查与选择性动作

RISA: Response Inspection and Selective Actions for Refusal Calibration in Large Language Models

Wenhan Chang, Tianqing Zhu, Ping Xiong, Shiyi Liao, Wanlei Zhou

arXiv 2609.00790首次发表:更新:

发表机构

School of Information Engineering, Zhongnan University of Economics and Law; City University of Macau(中南财经政法大学信息工程学院; 澳门城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出推理时框架RISA,通过检查初始响应并选择性纠正拒绝错误,在不更新大语言模型参数的情况下提升拒绝可靠性,同时保留模型效用。

AI 中文摘要

可靠的拒绝行为要求大语言模型(LLM)仅对良性提示作出回答,同时拒绝有害提示。错误的拒绝行为既可能使用户暴露于有害响应,也可能使用户无法获得有用答案。训练时对齐通过安全数据更新模型参数来改善拒绝行为,但需要额外的计算和训练成本。相比之下,推理时对齐旨在在推理过程中修改LLM的行为,而无需更新底层模型参数。现有的推理时方法主要依赖上下文安全提示、激活引导或解码控制。然而,大多数方法在未先确定初始响应是否恰当的情况下进行干预,可能会改变正确的拒绝或有用的答案。有效的选择性干预需要识别超出敏感关键词的提示意图,覆盖固定规则可能遗漏的语义变体,并使验证器适配不同的基础模型。为应对这些挑战,我们提出了响应检查与选择性动作(RISA),这是一种推理时框架,可检查初始响应并选择性纠正拒绝错误,无需更新基础模型。RISA首先使用固定上下文规则为明确案例分配拒绝分数;对于不匹配的案例,它使用校准后的线性探测从最后一层提示隐藏状态中推导拒绝分数。为适配不同的基础模型,RISA分别校准探测分数、表示支持边界和动作阈值。运行时,RISA将提示分数与初始拒绝状态结合,并应用动作策略仅在必要时进行干预。实验结果表明,RISA在提高拒绝可靠性的同时,很大程度上保留了模型效用,为LLM中感知响应的拒绝校准提供了实用解决方案。

英文摘要

Reliable refusal behavior requires Large Language Models (LLMs) to reject harmful prompts with only answering benign ones. Incorrect refusal behavior can either expose users to harmful responses or prevent users from obtaining useful answers. Training-time alignment improves refusal behavior by updating model parameters with safety data, but requires additional computation and training. In contrast, inference-time alignment aims to modify LLM behavior during inference without updating the underlying model parameters. Existing inference-time methods mainly rely on in-context safety prompting, activation steering, or decoding control. However, most of them intervene without first determining whether the initial response is already appropriate, potentially altering a correct refusal or a useful answer. Effective selective intervention therefore requires identifying prompt intent beyond sensitive keywords, covering semantic variations that fixed rules may miss, and adapting the verifier to different base models. To address these challenges, we propose Response Inspection and Selective Actions (RISA), an inference-time framework that inspects the initial response and selectively corrects refusal errors without updating the base model. RISA first uses fixed contextual rules to assign refusal scores to clear cases. For unmatched cases, it derives a refusal score from the final-layer prompt hidden state using a calibrated linear probe. To adapt to different base models, RISA separately calibrates the probe score, representation-support boundary, and action thresholds. At runtime, RISA combines the prompt score with the initial refusal status and applies an action policy to intervene only when necessary. Experimental results demonstrate that RISA improves refusal reliability while largely preserving model utility, offering a practical solution for response-aware refusal calibration in LLMs.

Comments17 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑