arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33298cs.LG

直接隐藏状态对齐:映射与控制LLM中的偏好表达

Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs

Fansheng Zhang, Shengran Guo, Zexiao Wang, Liang Yuan, Jiyuan Chen, Ruikun Luo

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出直接隐藏状态对齐(DHSA),通过残差竞争图映射偏好表达,并在推理时用少量参数干预隐藏状态,达到与DPO相当的对齐效果且可灵活启停。

中文摘要 AI 辅助

在许多场景中,后训练并不需要从零开始创建目标行为:基础模型已经能够产生该行为,但不够可靠。这将偏好对齐的部分任务从能力获取转变为行为表达。我们探究特定偏好在原生模型计算中如何表示,是什么阻止了支持目标的计算可靠地主导生成过程,以及这种结构能否直接指导控制。我们引入了残差竞争图(RCMs),它将行为偏好映射到原生残差计算的带符号因果效应上。在多个偏好域中,RCMs揭示了支持目标与竞争目标的效应共存、输入相关的组件角色,以及单个原生组件干预即可逆转偏好结果的情况。DPO大幅重组这些效应,并能削弱对立效应,但不保证将其消除。随后我们提出了直接隐藏状态对齐(DHSA),它将推理时的隐藏状态而非基础模型权重视为直接适应空间。RCM引导的因果激活状态转换(CAST)通过在少量偏好相关接口上进行局部状态干预来实现DHSA,同时冻结基础模型。仅使用256-16,384个控制器参数,CAST在三个偏好域中达到了与DPO竞争的性能水平,可以补充DPO训练的模型,并可在推理时启用或移除。

英文摘要

In many settings, post-training need not create the target behavior from scratch: the base model can already produce it, but not reliably. This shifts part of preference alignment from capability acquisition to behavioral expression. We ask how a specified preference is represented in native model computation, what prevents target-supporting computation from reliably dominating generation, and whether this structure can directly guide control. We introduce Residual Competition Maps (RCMs), which map a behavioral preference onto signed causal effects of native residual computation. Across preference domains, RCMs reveal coexisting target-supporting and target-competing effects, input-dependent component roles, and cases where a single native-component intervention reverses the preference outcome. DPO substantially reorganizes these effects and can weaken opposition without guaranteeing its removal. We then propose Direct Hidden-State Alignment (DHSA), which treats inference-time hidden states rather than base-model weights as the direct adaptation space. RCM-guided Causal Activation State Transition (CAST) implements DHSA through local state interventions at a small number of preference-relevant interfaces while freezing the base model. With only 256-16,384 controller parameters, CAST reaches DPO-competitive operating points across three preference domains, can complement DPO-trained models, and can be enabled or removed at inference time.

补充信息

↑