溢出感知的多值引导用于多元主义大语言模型对齐
Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment
浏览论文内容
中文总结 AI 辅助
针对多元主义对齐中多值引导的溢出问题,提出基于格拉姆矩阵的零成本修正方法,无需微调即可解耦方向,将净引导效应从+5.9%提升至+14.0%。
中文摘要 AI 辅助
激活引导通过在推理时向隐藏状态添加学习到的方向来控制大语言模型的行为,但现有方法一次只能处理一个概念。多元主义对齐要求不同利益相关者需要不同的价值侧重点,因此需要同时引导多个维度。我们表明,朴素引导会产生显著的溢出效应:针对一个价值意图的效果会泄漏到其他价值中。这与因果推断中的处理效应与溢出效应分解相呼应。我们将溢出追溯到引导方向的几何纠缠,由其格拉姆矩阵刻画,并从激活范数惩罚目标中推导出零成本修正,该修正精确地解耦每个方向的贡献。我们的端到端流程无需微调、无需奖励模型、无需手动提示工程:仅给定领域问题,它就能自动发现价值维度、提取方向、诊断纠缠并应用修正后的引导。在气候话语上,该修正将净引导效应从+5.9%提升至+14.0%,并在超过10万次成对判断上得到验证。
英文摘要
Activation steering controls LLM behavior at inference time by adding learned directions to hidden states, but existing methods handle one concept at a time. Pluralistic alignment, where different stakeholders need different value emphases, requires steering multiple dimensions simultaneously. We show that naive steering produces substantial spillover: the effect intended for one value leaks into others. This parallels the treatment-versus-spillover decomposition in causal inference. We trace spillover to geometric entanglement of steering directions, captured by their Gram matrix, and derive a zero-cost correction from an activation-norm-penalized objective that decouples each direction's contribution exactly. Our end-to-end pipeline requires no fine-tuning, no reward model, and no manual prompt engineering: given only domain questions, it automatically discovers value dimensions, extracts directions, diagnoses entanglement, and applies corrected steering. On climate discourse, the correction improves the net steering effect from +5.9% to +14.0%, validated over 100,000 pairwise judgments.
发表机构
- Stony Brook University(石溪大学)
机构由 AI 辅助整理,请以论文原文为准。