发表机构
Peking University; University of Chinese Academy of Sciences; Tsinghua University(北京大学; 中国科学院大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种在冻结残差状态上通过单向语义到价值通路实现低损伤LLM价值引导的方法,在保持语义的同时提升对齐并减少拒绝。
AI 中文摘要
价值引导应改变LLM的规范性优先级,同时保留其答案所依据的场景、事实和任务约束。传统的激活编辑往往同时改变两者。我们在冻结的残差状态上引入了一个可编辑的语义-价值接口,其中包含一个单向的语义到价值通路,该通路将价值识别建立在上下文之中。停止梯度通过该通路阻止反馈;交换一致性、主题去混杂和相关性的消除鼓励选择性编码。在推理时,编辑价值编码会产生残差增量,同时保持语义编码固定。在两个指令微调骨干网络上,该接口在相当的价值对齐水平下改善了语义保留并减少了良性拒绝。一种匹配的混合门控消融将表示学习与选择性编辑激活区分开来,维度匹配的探针证明了改进的编码选择性。与基于验证集选择的提示方法在LLaMA-3.1-8B上相比,该方法实现了相当的对齐(0.750对0.748)、更高的BERTScore(0.938对0.923)以及更少的矛盾(5.1%对7.6%)。人工评分和跨分类控制为低损伤价值引导提供了补充证据。
英文摘要
Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds value recognition in context. Stop-gradient blocks feedback through this pathway; swap consistency, topic de-confounding, and decorrelation encourage selective codes. At inference, editing the value code produces a residual delta while holding the semantic code fixed. On two instruction-tuned backbones, this interface improves semantic preservation and reduces benign refusals at comparable value alignment. A matched mixing-by-gating ablation separates representation learning from selective edit activation, and dimension-matched probes establish improved code selectivity. Against validation-selected prompting on LLaMA-3.1-8B, the method achieves comparable alignment (0.750 vs. 0.748), higher BERTScore (0.938 vs. 0.923), and fewer contradictions (5.1% vs. 7.6%). Human ratings and cross-taxonomy controls provide complementary evidence for low-damage value steering.