发表机构
BITS Pilani(比拉理工学院皮拉尼校区)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DiSCo框架,通过分布优先强制选择和四级情境梯度评估LLM文化偏好先验与可引导性,发现默认先验集中于英美且提示引导无法消除偏差。
AI 中文摘要
大型语言模型(LLMs)越来越多地部署在全球使用的助手中,然而它们在文化背景下的日常情境中的默认选择可能会系统性地偏向某些文化而非其他文化,影响本地化、用户信任和公平行为。现有的文化基准针对单一的“正确”答案评估准确性,当多种具有文化背景的回应都有效时,很难刻画LLM的文化偏好先验;它们还将默认偏好与情境驱动的适应混为一谈。我们提出DiSCo,一种分布优先的强制选择评估框架,该框架隔离默认文化先验,并通过四级情境梯度(C0--C3)测试可引导性。使用源自BLEnD的DiSCo-Bench(304个条目),涵盖12种文化,我们评估了六个不同的指令微调LLM。默认先验高度集中,英国和美国合计吸收了约35%的所有选择,尽管它们仅代表12种文化中的2种。最关键的是,基于提示的引导持续扩大高资源与低资源文化之间的选择差距,而注入明确的文化事实产生的分布扰动可忽略不计,证实文化偏好偏差无法仅通过基于提示的个性化来解决。
英文摘要
Large language models (LLMs) are increasingly deployed in globally used assistants, yet their default choices in culturally grounded everyday situations can systematically favour some cultures over others, affecting localisation, user trust, and equitable behaviour. Existing cultural benchmarks evaluate accuracy against a single "correct" answer, making it difficult to characterise an LLM's cultural preference prior when multiple culturally grounded responses are all valid; they also conflate default preferences with context-driven adaptation. We propose DiSCo, a distribution-first forced-choice evaluation framework that isolates default cultural priors and tests steerability via a four-level context gradient (C0--C3). Using DiSCo-Bench (304 items) derived from BLEnD spanning 12 cultures, we evaluate six diverse instruction-tuned LLMs. Default priors are heavily concentrated, with UK and US together absorbing approximately 35\% of all selections despite representing only 2 of 12 cultures. Most critically, prompt-based steering consistently widens the selection gap between high- and low-resource cultures, and injecting explicit cultural facts produces negligible distributional disruption, confirming that cultural preference bias cannot be resolved through prompt-based personalisation alone.