发表机构
School of Computer Science and Technology, Tianjin University; Singapore University of Technology and Design(天津大学计算机科学与技术学院; 新加坡科技设计大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出协同评估框架分析大语言模型的文化对齐与多样性权衡,发现追求文化对齐会导致严重的文化扁平化,揭示该现象源于神经网络优化的低秩偏差,为后训练范式改进提供方向。
AI 中文摘要
文化微调已成为构建具有文化感知能力的大语言模型(LLM)的事实上的范式,然而现有仅针对对齐分数的优化,通过系统性地掩盖内在的文化多样性,仅呈现了文化保真度的不完整图景。这种单维度的评估视角引出了一个根本问题:模型是真正感知到不同的文化细微差别,还是仅仅记住了主导的文化价值观?为解决该问题,我们提出了一个协同评估框架,该框架同时将文化对齐与多样性形式化。通过在世界价值观调查(World Values Survey)上对六个主流大语言模型(LLM)进行广泛的基准测试,该框架揭示了一种系统性且关键的权衡:追求文化对齐始终会带来多样性的严重损失,导致出现严重的“文化扁平化”现象。对这种行为转变进行调查后,我们证明这些表面的对齐收益源于模型人为地锚定到占主导地位的多数群体,收敛到单一的响应模式,该模式消除了人类群体固有的异质分布。至关重要的是,我们的机制分析表明,这种多样性崩溃不仅是一种行为异常,更可能是神经网络优化中固有低秩偏差的结构性后果。因此,我们的发现揭示了当前后训练范式的局限性,并呼吁转向能够保留跨文化多元性的对齐目标。
英文摘要
Cultural fine-tuning has become the de facto paradigm for building culture-aware large language models (LLMs), yet existing optimization exclusively for alignment scores provides an incomplete portrait of cultural fidelity by systematically obscuring inherent cultural diversity. This unidimensional evaluation lens prompts a fundamental question: do models genuinely perceive distinct cultural nuances, or do they merely memorize dominant cultural values? To address this, we propose a synergistic evaluation framework that jointly formalizes cultural alignment and diversity. Through extensive benchmarking of six mainstream LLMs on the World Values Survey, this framework uncovers a systematic and critical trade-off: the pursuit of cultural alignment consistently incurs an acute expense of diversity, leading to severe "cultural flattening." Investigating this behavioral shift, we demonstrate that these superficial alignment gains stem from models artificially anchoring to dominant majorities, converging onto a monolithic response pattern that wipes out the heterogeneous distributions inherent to human groups. Crucially, our mechanistic analysis suggests that this diversity collapse is not merely a behavioral anomaly but more likely a structural consequence of the low-rank bias inherent in neural network optimization. Therefore, our findings expose the limitations of current post-training paradigms and call for a shift toward alignment objectives that preserve cross-cultural pluralism.
CommentsAccepted at EMNLP 2026 (Findings)