arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为谁保障安全?面向受控大语言模型安全拒绝的边界感知自蒸馏

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

Alejo López-Ávila, Iker García-Ferrero, Jezabel Garcia, Antonio Tiene, Román Orús

arXiv 2609.04482首次发表:更新:

发表机构

Multiverse Computing(Multiverse Computing)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型安全对齐的窄边界需求,提出结合受控主题生成等的离线自蒸馏框架,实验显示其可提升目标域拒绝率并降低不安全响应率,同时需权衡安全与可用性。

AI 中文摘要

安全对齐通常被表述为主题层面的问题:这个主题是否有害?而实际部署中需要的是更狭窄的问题。公民教育辅导工具和公共部门助手可能共享同一个基础模型,但在同一主题内需要不同的边界,即拒绝有针对性的政治操纵,同时仍回答关于同一选举的事实性问题。我们将此定义为窄边界安全,并引入一种离线自生成框架,该框架结合受控主题生成、覆盖修复、分布内补偿数据以及有害-良性对,用于训练和评估。单次生成会导致19.88%的提示词没有被接受的拒绝痕迹,而升级重试则仅留下0.20%的提示词未被拒绝。在使用Qwen3-8B进行的政治说服实验中,基于通过Escalate完成的拒绝数据进行训练,使目标域拒绝率从9.47%提升至84.75%,并将三个更广泛危害性基准的平均不安全响应率从26.26%降至0.14%,但使XSTest过度拒绝率从2.00%升至74.00%。在另一项匹配对比中,用经验证的目标模型响应替代外部响应,使过度拒绝率从15.20%降至5.20%。边界对数据使保留对的顺从侧过度拒绝率从32.94%降至4.16%,而有害侧拒绝率仅从91.88%降至87.72%。这些结果表明,数据构成可控制安全与可用性之间的权衡,且安全对齐应在预期拒绝边界的两侧进行评估。

英文摘要

Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A civics tutor and a public-sector assistant may share a base model yet need different boundaries inside the same topic, refusing targeted political manipulation while still answering factual questions about the same election. We formulate this as narrow-boundary safety and introduce an offline self-generated framework combining controlled topic generation, coverage repair, in-distribution compensation data, and harmful-benign pairs for training and evaluation. Single-shot generation leaves 19.88% of prompts without accepted refusal traces, whereas escalating retries leave 0.20%. On political persuasion with Qwen3-8B, training on refusal data completed through Escalate increases target-domain refusal from 9.47% to 84.75% and reduces the mean unsafe-response rate across three broader harmfulness benchmarks from 26.26% to 0.14%, but increases XSTest over-refusal from 2.00% to 74.00%. In a separate matched comparison, replacing external responses with verified target-model responses reduces over-refusal from 15.20% to 5.20%. Boundary-pair data reduces comply-side over-refusal on held-out pairs from 32.94% to 4.16%, while harmful-side refusal decreases only from 91.88% to 87.72%. These results show that data composition controls the safety and usability trade-off, and that safety alignment should be evaluated on both sides of the intended refusal boundary.

Comments22 pages, 14 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑