发表机构
University of British Columbia; Vector Institute for AI; Queen’s University(不列颠哥伦比亚大学; 向量人工智能研究所; 女王大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过引入26K样本基准,验证了LLM转向向量几何结构与人类价值理论的一致性,发现分布驱动方法优于行为中心方法,且几何保真度随模型规模提升但受指令微调影响。
AI 中文摘要
随着大型语言模型(LLMs)越来越多地部署在对齐敏感的环境中,激活转向(activation steering)已成为一种轻量级的推理时替代微调方法(如RLHF、DPO)进行行为控制的手段。然而,现有工作通常针对孤立行为验证转向,尚不清楚转向向量是否编码了连贯的语义结构,还是仅仅利用了特定行为的捷径。我们研究了LLM转向向量的潜在几何结构是否反映了人类价值观和道德中理论规定的结构。以Schwartz的基本人类价值理论作为主要细粒度框架,我们引入了一个涵盖20种人类价值观的26K样本基准,并分析了多种模型系列和规模下的分布驱动方法(如CAA、SphericalSteer、ODESteer)和以行为为中心的方法(如COLD-Steer、BiPO)。我们发现,分布驱动方法恢复的人类价值拓扑与理论预测一致(Spearman ρ高达0.51,p < 10^-13)。相比之下,以行为为中心的方法实现了相当的转向性能,但与预期价值几何的相关性很小。几何保真度随模型规模提高而改善,但在指令微调后下降。最后,更好的几何对齐也导致跨价值观的更符合人类一致性的迁移:正确转向一个价值观会提升兼容的价值观并抑制对立的价值观。代码和数据可在以下网址获取:此https URL。
英文摘要
As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically validates steering on isolated behaviors, leaving it unclear whether steering vectors encode coherent semantic structure or merely exploit behavior-specific shortcuts. We investigate whether the latent geometry of LLM steering vectors reflects theory-specified structure in human values and morality. Using Schwartz's Theory of Basic Human Values as our primary fine-grained framework, we introduce a 26K-sample benchmark covering 20 human values and analyze distribution-driven methods (e.g., CAA, SphericalSteer, ODESteer) and behavior-centric approaches (e.g., COLD-Steer, BiPO) across diverse model families and sizes. We find that distribution-driven methods recover human value topologies aligned with theoretical predictions (Spearman $ρ$ up to 0.51, $p < 10^{-13}$). In contrast, behavior-centric methods achieve comparable steering performance but show little correlation with the expected value geometry. Geometric fidelity improves with model scale but drops after instruction tuning. Finally, better geometric alignment also leads to more human-consistent transfer across values: steering one value correctly lifts compatible values and suppresses opposing ones. Code and data are available at: https://github.com/DeepRCL/Steering_Geometry.
CommentsAccepted to EMNLP 2026 Main Conference (top 15.4%)