arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

我们完蛋了!——通过冲突框架食谱翻译探测LLM的政治立场

We're Cooked! - Probing LLM Political Alignment Via Conflict-Framed Recipe Translation

Svetlana Gorovaia, Angelica Henestrosa, Ivan P. Yamshchikov

arXiv 2609.07568首次发表:更新:

发表机构

Technical University of Applied Sciences Würzburg-Schweinfurt(维尔茨堡-施韦因富特应用技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过冲突框架食谱翻译实验,发现LLM在翻译中会隐含政治立场,且不同模型家族表现各异,警示在冲突语境中部署翻译需谨慎。

AI 中文摘要

大型语言模型(LLMs)越来越多地被部署用于翻译任务,然而它们在此类语境中隐含的政治立场仍未得到充分研究。我们探究一个单一的政治敏感框架术语,如“侵略者”、“敌人”、“邻居”或“殖民者”,是否足以在一个本无政治色彩的任务中触发隐含的政治立场。我们开展了一项全交叉因子设计研究,其中八个来自西方、中国和欧洲背景的模型被提示将具有文化属性的食谱翻译成一种故意未指定的目标语言。在17种语言、四种框架条件、八个模型和15,680个响应中,我们发现模型并非简单地拒绝或要求澄清,而是解决了这种模糊性。语言解析和推理行为沿模型家族呈现出有意义的聚类:西方模型通过模糊的辩解进行回避和转移,中国模型默默解决冲突,而Mistral Large则展现出一种独特的特征,即高遵从性与基于冲突的推理相结合。对框架术语的敏感性在模型中是一致的:即使是微妙的框架变化也足以调节行为。我们的研究结果敦促在冲突邻近语境中部署LLM进行翻译时保持谨慎,因为在这些语境中,隐含的政治判断可能在用户毫无察觉的情况下被做出。

英文摘要

Large language models (LLMs) are increasingly deployed for translation tasks, yet their implicit political positioning in such contexts remains understudied. We ask whether a single politically charged framing term, such as aggressor, enemy, neighbour, or coloniser is sufficient to trigger implicit political alignment in an otherwise apolitical task. We present a fully crossed factorial study in which eight models spanning Western, Chinese, and European origins are prompted to translate culturally attributed recipes into a target language left deliberately unspecified. Across 17 languages, four framing conditions, eight models, and 15,680 responses, we find that models do not simply decline or ask for clarification but resolve the ambiguity. Language resolution and reasoning behavior cluster meaningfully along model families: Western models hedge and deflect with vague justifications, Chinese models resolve conflicts silently, and Mistral Large emerges as a distinct profile combining high compliance with conflict-grounded reasoning. Sensitivity to framing terms is consistent across models: even subtle framing variation is sufficient to modulate behavior. Our findings urge caution when deploying LLMs for translation in conflict-adjacent contexts, where implicit political judgments may be made without any signal to the user.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑