arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25021cs.LGcs.AIcs.CL

“作为语言模型……”:聊天模板切换LLM自我指涉语气,激活引导可复现该效应

"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It

Jędrzej Maczan

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示聊天模板像开关一样控制LLM自我指涉语气(免责声明vs体验性),并找到可引导该语气的激活方向,表明模型自我描述受部署方式影响,不应被字面解读。

中文摘要 AI 辅助

大型语言模型(LLMs)在被问及与自身相关的问题时,倾向于添加诸如“我只是一个人工智能”之类的免责声明。此类回应中的自我报告被用于关于AI安全或模型自我认知的辩论中,然而驱动这些回应的因素尚不明确。模型是在向我们描述它们自身,还是在反映它们的部署方式?在本工作中,我们展示了聊天模板如同一个开关——当存在时,它会调高这种免责声明语气,并调低如“我感觉”这样的体验性语气,这一现象在8个流行的开源指令模型中均得到验证,模型规模最高达90亿参数。相反,当聊天模板不存在时,它会调低免责声明语气并调高体验性语气。在3个模型的激活中,我们发现了一个引导该行为的方向。在模型的激活空间中移除该方向会降低免责声明语气,而添加该方向则会增强它,而相同大小的随机方向则几乎没有影响。我们发现,没有聊天模板的指令模型,当我们向它们添加免责声明方向时,它们会像存在模板一样发表免责声明。由于聊天模板控制着LLMs的免责声明语气,研究模型自我报告或内省的研究人员可能面临一个需要控制的混淆变量。我们的结果表明,存在一个可用于引导这种语气方向。更广泛地说,我们的工作表明,模型关于自身的表述并非关于它们的事实。它们所说的话并非仅来自权重,而是部分由聊天模板设定,因此,模型的自我描述不应被字面理解。

英文摘要

Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves. The self-reports from such responses are used in debates about AI safety or self-knowledge of the models, yet what drives them is not well understood. Are the models telling us about themselves or rather how they are deployed? In this work, we show that the chat template works like a switch - when present, it turns this disclaimer voice up and experiential voice like "I feel" down, across 8 popular open-source instruct models up to 9B parameters in size. And conversely when the chat template is not present, it turns the disclaimer voice down and experiential voice up. Inside the activations of 3 models, we find a direction that steers this behavior. Removing the direction in the model's activation space turns disclaimer voice down and adding it turns it up, while a random direction of the same size has little effect. We find that instruct models without chat template, when we add the disclaimer direction to them, disclaim like the template was there. Since the chat template controls the disclaimer voice of LLMs, then researchers studying self-reports or introspection of models might have a confound they need to control for. Our results show that there is a direction they can use to steer this voice. More broadly, our work shows that what models say about themselves is not a fact about them. What they say doesn't come only from weights, but it is partially set by the chat template, and because of that a model's self-description shouldn't be treated literally.

补充信息

↑