发表机构
DataSpike; Metropolia University of Applied Sciences; Neapolis University Pafos(DataSpike; 梅托波利亞應用科學大學; 帕福斯尼阿波利斯大學)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究分析18个开源LLM的神话知识表征,发现文化表征存在于模型残差流中,但解码器因提示语言偏差无法解码非主导文化内容,还发布了相关跨文化研究工具与基准。
AI 中文摘要
开源大语言模型(LLM)能可靠地命名宙斯、朱庇特和托尔,但在代表性较低的传统神话(如芬兰、斯拉夫、埃及或中国神话)中,对应神祇的一致性要差得多。我们探究这种文化默认偏差在模型内部的产生位置。基于汤普森主题实体的跨文化平行子结构,我们对来自8个架构家族的18个开源LLM,采用线性探测、logit lens(对数透镜)、激活修补和输出提取方法进行分析。残差流能清晰区分不同文化,远高于名称字符串基线,但解码器会将特定文化的令牌映射到主导文化的令牌上。失败出在读取阶段,而非表征阶段。用目标文化的母语与英语提问产生的失败,在语言内聚类但跨语言解耦:解码器受提示语言控制。我们发布了基于每个实体的(探测、输出)分解框架、基于引用锚定的跨文化基准、针对语言条件读取的内部与跨模式相关性测试,以及18个模型的每个实体预测结果。
英文摘要
Open-source LLMs reliably name Zeus, Jupiter, and Thor, but recover their counterparts in less-represented traditions like Finnish, Slavic, Egyptian, or Chinese mythology far less consistently. We ask where inside the model this cultural default is produced. On a parallel cross-cultural substrate of Thompson-motif entities, we instrument 18 open-source LLMs from 8 architecture families with linear probing, logit lens, activation patching, and output extraction. The residual stream cleanly distinguishes cultures, well above a name-string baseline, yet the decoder collapses culturally-specific tokens onto dominant-tradition ones. The failure is at readout, not at representation. Asking the same question in the target culture's native language versus English produces failures that cluster within language but decouple across language: the decoder is gated on prompt language. We release a per-entity (probe, output) decomposition framework, a citation-anchored cross-cultural ground truth, a within- versus cross-mode correlation test for language-conditioned readout, and per-entity predictions for all 18 models.
Comments45 pages, 23 figures, 18 tables. Dataset: https://huggingface.co/datasets/Aragoner/folkmotif Code: https://github.com/AragonerUA/folkmotif