arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34771cs.AIcs.CLcs.LG

模型内部何时有用?探索表示工程在LLM安全中的作用

When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety

Tianyi Guan, Jianhui Chen, Liangming Pan

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过匹配评估发现,表示工程在安全控制上仅在低数据场景有优势,在监控上成本低但精度略逊,总体不替代行为防护,但可互补提升LLM安全性。

中文摘要 AI 辅助

可靠的AI安全防护需要两类机制:一类是减少不安全行为的控制机制,另一类是在模型交互过程中检测安全风险的监控机制。已建立的行为安全防护包括优化模型输出的对齐方法和评估交互文本的文本监控器。表示工程则通过读取或修改模型内部状态来发挥作用,但这些方法之间的相对优势尚不明确,因为它们通常在不同的设置下进行评估。我们提出了一个跨两条轨道的匹配评估。对于安全控制,我们比较了DPO(一种行为对齐方法)与三种表示引导方法在鲁棒性、实用性和细粒度方面的表现。DPO提供了最强的整体控制,并且通常随着训练数据的增加而改善,尽管其安全性在后续的良性微调后可能会下降。表示引导主要在低数据设置下具有竞争力,尤其是使用高质量对比数据时。对于安全监控,我们比较了表示探针与微调的和开放权重的文本监控器在全响应检测、早期检测和计算成本方面的表现。专门的文本监控器实现了最强的整体检测准确性,而表示探针在边际成本显著更低的情况下仍具有竞争力。最后,监控器引导的干预措施恢复了DPO在良性微调后损失的大部分安全性,且几乎没有额外的过度拒绝。总体而言,表示工程通常不能替代行为安全防护,但在特定条件下提供了实际优势,并能带来互补的安全效益。

英文摘要

Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect safety risks during model interactions. Established behavioral safeguards include alignment methods that optimize model outputs and text monitors that assess interaction text. Representation engineering instead reads or modifies internal model states, but the relative strengths of these approaches remain unclear because they are often evaluated under different settings. We present a matched evaluation across two tracks. For safety control, we compare DPO, a behavioral alignment method, with three representation steering methods across robustness, practicality, and granularity. DPO provides the strongest overall control and generally improves with increasing training data, although its safety can degrade after subsequent benign fine-tuning. Representation steering remains competitive primarily in low-data settings, particularly with high-quality contrastive data. For safety monitoring, we compare representation probes with fine-tuned and open-weight text monitors across full-response detection, early detection, and computational cost. Specialized text monitors achieve the strongest overall detection accuracy, while representation probes remain competitive at substantially lower marginal cost. Finally, monitor-guided interventions recover much of the safety lost by DPO after benign fine-tuning, with little additional over-refusal. Overall, representation engineering does not generally replace behavioral safeguards, but offers practical advantages under specific conditions and can provide complementary safety benefits.

发表机构

  • State Key Laboratory of Multimedia Information Processing, Peking University(北京大学多媒体信息处理国家重点实验室)
  • School of Computer Science, Peking University(北京大学计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

↑