arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型中的文化错位:通过针对性微调进行检测、测量与缓解

Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning

Antoni Czolgowski, Abel Iyasele

arXiv 2609.04485首次发表:更新:

发表机构

University of Colorado Boulder(科罗拉多大学博尔德分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究以三个开放权重LLM为对象,量化其跨文化错位,发现针对性LoRA微调可降低偏差但仅重新分配偏差,是首个针对最坏情况角色用LoRA微调缓解跨文化偏差的研究。

AI 中文摘要

我们针对来自美国的Gemma3-12B、波兰的Bielik-11B-v3、中国的Qwen3-4B这三个开放权重大语言模型(LLM),采用世界价值观调查第七波针对三个国家的63个人口统计角色的数据,使用归一化Wasserstein距离量化分布错位。与预期相反,没有模型偏向其母国:中国开发的Qwen3-4B在其本国中国人口上表现最差(W1=0.436,是整个模型-国家矩阵中最高的错位值)。对5个最坏情况角色进行针对性LoRA微调,仅需不到1200个训练对,在单个GPU上耗时不到15分钟,可使Bielik-11B的偏差降低16.8%(p_Bonf=0.002,d=-4.4),且所有5个目标角色均得到改善。然而,国家层面分解显示,微调是重新分配而非消除偏差:Bielik的最坏情况角色完全从美国老年人转变为中国老年人,校正前后的集合无重叠。据我们所知,这是首个针对最坏情况人口统计角色采用LoRA微调缓解跨文化偏差的研究。

英文摘要

We evaluate three open-weight LLMs (Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from China) against World Values Survey Wave 7 data for 63 demographic personas across three countries, using normalized Wasserstein distance to quantify distributional misalignment. Contrary to expectations, no model favors its home country: the Chinese-built Qwen3-4B performs worst on its own Chinese population (W1 = 0.436, the highest misalignment in the entire model x country matrix). Targeted LoRA fine-tuning on the five worst-case personas, requiring fewer than 1,200 training pairs and under 15 minutes on a single GPU, reduces bias by 16.8% for Bielik-11B (p_Bonf = 0.002, d = -4.4) with all five targets improving. However, country-level decomposition reveals that fine-tuning redistributes rather than removes bias: Bielik's worst-case personas swap entirely from American to Chinese elderly, with zero overlap between pre- and post-correction sets. To our knowledge, this is the first study to target worst-case demographic personas with LoRA fine-tuning for cross-cultural bias mitigation.

Comments33 pages, 14 figures. Extended version of a paper published in the proceedings of OSSConf 2026, Zilina, Slovakia. Code and data: https://github.com/AntoniCzolgowski/llm-cultural-bias

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑