arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22224cs.CLcs.LG

从特征向量到电路:追踪语言模型中的拒答与谄媚行为

From Trait Vectors to Circuits: Tracing Refusal and Sycophancy Through Language Models

  • University of Amsterdam(阿姆斯特丹大学)
  • Safe AI Netherlands(荷兰安全人工智能)

机构由 AI 辅助整理,请以论文原文为准。

Oscar Miró López-Feliu, Maya Ozbayoglu

AI总结:

本研究通过将特征向量分割为重建和传输电路,测试Qwen2.5-7B-Instruct中拒答与谄媚行为的机制,发现引导干预的电路并非行为产生的电路,拒答电路紧凑忠实,谄媚重建部分忠实。

AI中文摘要:

激活空间中的一个方向,当被引导时能改变与安全相关的行为,但这并不一定是模型自身用于产生该行为的方向。因此,我们探究引导是否通过未修改模型的计算起作用,还是通过一组不同的组件起作用,研究了先前工作中已提取并验证的两个特征的方向:Qwen2.5-7B-Instruct中的拒答和谄媚。对于每个特征,我们使用特征向量将计算分为向量之前的重建电路和向量之后的传输电路。然后,我们测试恢复该坐标是否返回被消融移除的行为。对于拒答,两个电路都紧凑且忠实,单独恢复该坐标几乎能恢复因消融而丢失的全部拒答信号,并且围绕该向量构建的电路在约一半的边数下匹配直接输入到输出电路的忠实度。对于谄媚,传输电路紧凑,但重建电路更广泛且仅部分忠实。因此,传输引导干预的电路并不自动是产生该行为的电路,我们相应地报告每个电路的反事实、目标和行为测试。

英文摘要:

A direction in activation space that changes safety-relevant behavior when steered is not necessarily one the model uses to produce that behavior on its own. We therefore ask whether steering acts through the computation of the unmodified model or through a different set of components, studying two traits whose directions have been extracted and validated in prior work: refusal and sycophancy in Qwen2.5-7B-Instruct. For each, we use the trait vector to split the computation into a reconstruction circuit before the vector and a transmission circuit after it. We then test whether restoring the coordinate returns behavior removed by ablation. For refusal, both circuits are compact and faithful, restoring the coordinate alone recovers almost all of the refusal signal lost to ablation, and a circuit built around the vector matches the faithfulness of a direct input-to-output circuit at roughly half the edges. For sycophancy, transmission is compact, but reconstruction is broader and only partially faithful. Circuits that transmit a steering intervention are therefore not automatically the circuits that produce the behavior, and we report the counterfactual, target, and behavioral test for every circuit accordingly.

↑