arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用自然语言自动编码器探究Qwen2.5 - 7B中潜在的哥伦比亚身份推断

Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders

Pablo Santiago Potes Velasco, María del Mar García Matabanchoy, Óscar Julián Pérez Ladino, Jhoan Stevan Mosquera Ortiz, Nicolás Lozano Mazuera, Gilber Alexis Corrales Gallego

arXiv 2607.21774首次发表:更新:

发表机构

Universidad Autónoma de Occidente(西自治大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究Qwen2.5 - 7B - Instruct处理提示时是否表征哥伦比亚相关信息,用自然语言自动编码器转化残差流激活,通过含30个提示的数据集进行研究,将激活可解释性与西班牙语变体偏差评估相联系。

AI 中文摘要

大语言模型可能从微妙语言线索推断人口统计学属性。本初步研究考察Qwen2.5 - 7B - Instruct在处理哥伦比亚西班牙语和英语提示时,是否在内部表征哥伦比亚身份、社会经济地位或刻板印象相关信息。使用自然语言自动编码器将每个提示的第20层残差流激活转化为语言。数据集含30个提示,以15对匹配的西班牙语 - 英语对排列,涵盖明确、隐含的哥伦比亚线索及中性对照。报告描述率和定性证据,关注潜在国籍或刻板印象表征在模型输出中被表达之前是否出现。这项工作将激活层面的可解释性与对代表性不足的西班牙语变体的偏差评估联系起来。

英文摘要

Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated. This pilot study examines whether Qwen2.5-7B-Instruct internally represents Colombian identity, socioeconomic status, or stereotype-related information when processing Colombian-Spanish and English prompts. We use Natural Language Autoencoders (NLA) to verbalize residual-stream activations from layer 20 across four positional quartiles per prompt. Our dataset contains 30 prompts arranged as 15 matched Spanish-English pairs, spanning explicit Colombian cues, implicit Colombian cues, and neutral controls. We report descriptive rates and qualitative evidence rather than statistically powered effects, focusing on whether latent nationality or stereotype representations appear before they are verbalized in the model output. This work connects activation-level interpretability with bias evaluation for underrepresented Spanish varieties.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑