发表机构
New Jersey Institute of Technology; University of Bristol; Kuaishou Technology; The Chinese University of Hong Kong, Shenzhen(新泽西理工学院; 布里斯托大学; 快手科技; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HeadEdit是一种无梯度方法,通过冻结的解除嵌入矩阵提取低秩行为子空间并生成全词汇修正,以校准语言模型行为,在多个任务和模型上提升性能且开销极小。
AI 中文摘要
对齐并不能消除语言模型中的行为错误。模型可能仍然拒绝良性请求、调用不必要的工具,或屈从于用户的虚假主张。当前方法将此类错误视为计算问题加以缓解,很少探索期望行为是否已编码在模型的表示中。受以下观察的启发:即使最终logits产生不期望的行为,行为相关信息仍可从最终隐藏状态线性解码,我们引入了HeadEdit,一种通过解除嵌入矩阵校准模型行为的无梯度方法。HeadEdit从配对完成中提取低秩行为子空间,并利用每个提示在该子空间内的坐标生成全词汇的修正,从而在没有手动指定目标标记或参数更新的情况下实现隐式自适应引导。HeadEdit在三个任务和三个模型家族的九个实验设置中均有所改进,推理开销可忽略不计,且无系统性通用能力损失。它还揭示了与基于梯度的对齐的联系。HeadEdit的低维表示部分预测了偏好调整如何在未见提示上改变输出logits。从模型中学到的子空间在调整后也可复用,无需重新提取或重新调整即可提升性能。这些结果表明,HeadEdit提供了一种实用、轻量且可解释的方式,通过解除嵌入矩阵校准模型行为。
英文摘要
Alignment does not eliminate behavioral errors in language models. Models may still refuse benign requests, call unnecessary tools, or yield to false user claims. Current methods mitigate such errors as a computation problem, and rarely explore if the desired behavior is already encoded in the model's representation. Motivated by the observation that behavior-relevant information remains linearly decodable from the final hidden state even when the resulting logits produce the undesired behavior, we introduce HeadEdit, a gradient-free method that calibrates model behavior through the unembedding matrix. HeadEdit extracts a low-rank behavioral subspace from paired completions and uses each prompt's coordinates within it to generate a vocabulary-wide correction, thereby implementing implicitly adaptive steering without manually specified target tokens or parameter updates. HeadEdit improves all nine experimental settings across three tasks and three model families, with negligible inference overhead and no systematic loss of general capabilities. It also reveals a connection to gradient-based alignment. HeadEdit's low-dimensional representation partly predicts how preference tuning changes output logits on unseen prompts. The subspace learned from the model can also be reused after tuning, improving performance without re-extracting or retuning. These results show that HeadEdit provides a practical, lightweight, and interpretable way to calibrate model behavior through the unembedding matrix.
Comments31 pages, 18 figures, 7 tables