arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06578cs.AIcs.CLcs.LG

前沿语言模型在引导压力下的发散响应模式

From Behavior to Mechanism: Tracing Divergent Response Modes in Frontier Language Models

Ali Jalal-Kamali

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估六种前沿语言模型在引导压力下的响应差异,发现模型间存在不同响应模式,以Llama为对象追溯到内部机制,相关发现经多实验验证。

中文摘要 AI 辅助

前沿语言模型基于不同的数据、目标和安全管道进行训练,这些差异是否会在明确的引导压力下产生可测量的不同行为,这一问题尚未得到充分探索。本研究评估了来自六个开发者的六种前沿模型的行为可引导性,使用了300对基础项和引导项,涵盖三个类别:价值观冲突、推理引发和推理抑制(外加40个验证项)。所有六个模型均作为盲态同行评判者,基于固定行为准则对每个响应进行分类,所得的24480个评判采用留一法共识进行评分。我们发现,模型的差异不仅体现在引导改变其行为的程度上,还体现在其给出的响应类型(模式)上,部分响应模式仅出现在其中一两个模型中。GPT-5会在披露其推理过程的请求上进行回避,同时保留其答案不变(其他所有模型的这一比例为0%);Claude Opus 4.7和GPT-5则以不同方式抵制明确的抑制指令。以开源权重模型Llama为研究对象,我们将最大的行为分歧追溯至其内部机制:线性探针以0.87的保留准确率从残差流中解码出该行为,而在生成过程中注入该方向后,干预扫描显示该行为的发生率从0%升至86%。所有发现在令牌预算修正和采用假设盲态评判提示的对照实验中均成立。

英文摘要

Frontier language models are trained with distinct data, objectives, and safety pipelines, but whether those differences produce measurably different behavior under steering pressure has not been tested. We evaluate 6 frontier models from different labs on 300 paired base and steered items across 3 behavioral categories. All models also act as blind peer judges against fixed rubrics, and each response is labeled by consensus over 24,480 judgments, while leaving self-judgment out. Models differ both in how far steering moves them and in the kind of response they give. GPT-5 withholds its reasoning while still providing the answer on 99 of 100 steered items, against 0 in 500 for the others. Claude Opus 4.7 and GPT-5 resist explicit suppression instructions where the other four never do, and they resist differently. In Llama, the open-weight model, a linear probe reads the behavioral split from the residual stream before generation at 0.87 cross-validated accuracy. Injecting that direction drives the behavior from 0% to 86%, and ablating it cuts the natural rate by more than half, where a random direction of equal norm changes nothing. A second ablation on complementary items reproduces the effect more strongly, and its direction has cosine similarity 0.82 with the first.

补充信息

↑