arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过因果追踪理解语言实现选择对大语言模型立场的影响

Understanding the Impact of Linguistic Realization Choices on LLM Stance with Causal Tracing

Langchen Huang, Sebastian Padó, Franziska Weeber

arXiv 2607.20115首次发表:更新:

发表机构

Institute for Natural Language Processing, University of Stuttgart(斯图加特大学自然语言处理研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究语言结构对大语言模型立场的影响,以政治立场判断为案例,扩展数据集得六种重写类型,通过对四个模型实验及激活修补发现,中晚期解码器层尤其是最终提示位置的块输出,对恢复原立场分布有最强信号。

AI 中文摘要

大语言模型(LLMs)对提示和输入表述敏感。现有研究聚焦词汇实现,忽视结构选择。本文研究语言结构能否系统改变LLM决策及在模型内因果定位。以政治立场判断为例,扩展英语政治声明数据集,得六种语言重写类型。对四个开放权重模型实验表明立场不稳定影响意义保留和反转重写。因输出变化揭示重写影响立场但不知模型内位置,故用激活修补,结果显示中晚期解码器层,尤其最终提示位置的块输出,提供最强恢复信号。

英文摘要

Large language models (LLMs) are known to be sensitive to prompt and input formulations. However, existing studies have focused on lexical realization and largely ignored constructional choice. This paper studies whether linguistic construction can systematically shift LLM decisions and where these shifts can be causally localized inside the model. We use political stance judgment as a meaning-sensitive case study and extend an English political statements dataset, resulting in six controlled linguistic rewrite types that preserve or invert the meaning of a statement. Experiments on four open-weight models show that stance instability affect both meaning-preserving and meaning-inversing rewrites. Because output shifts reveal that rewrites affect stance, but not where in the model, we apply activation patching, where activations from the original statement are substituted into the forward pass for the rewritten statement and measure which components recover the original stance distribution. The results show that mid-to-late decoder layers, especially block outputs at the final prompt position, provide the strongest restoration signal.

CommentsKONVENS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑