大语言模型混淆了哪些价值观?基于施瓦茨的识别研究
Which Values Do LLMs Confuse? A Schwartz-Based Recognition Study
浏览论文内容
中文总结 AI 辅助
研究大语言模型对施瓦茨十个基本价值观的识别,通过俄语情境文本评估 21 个指令微调的模型运行,分析准确率、混淆情况等,推动结合精确准确率、排序恢复和有向错误分析的价值识别评估。
中文摘要 AI 辅助
大语言模型越来越多地通过它们认可的价值观来评估,但这种评估预先假定模型能够识别具体情境中表达的价值观。我们将此前提研究为对施瓦茨的十个基本价值观的受控 top-1 识别。我们的评估集包含 1000 篇俄语情境文本,在十个价值观上保持平衡,每个项目由两名人类注释者独立标注。我们在固定的排序响应协议下评估 21 个指令微调的大语言模型运行;20 个具有可靠输出的运行形成语义面板。汇总的 Acc@1 为 0.683,Acc@3 为 0.892,表明模型在定位正确动机区域时,对相近替代选项的排序不稳定。相邻价值观占语义错误的 50.9%,而特定检查点的空值下为 24.4%。八个有向混淆在检查点和人类确认的子集中反复出现。结果推动了结合精确准确率、排序恢复和有向错误分析的价值识别评估。
英文摘要
Large language models are increasingly evaluated through the values they endorse, but such evaluations presuppose that models can identify the value expressed in a concrete situation. We study this prerequisite as controlled top-1 recognition over Schwartz's ten basic values. Our evaluation set contains 1,000 Russian situational texts, balanced across the ten values and independently labeled by two human annotators per item. We evaluate 21 instruction-tuned LLM runs under a fixed ranked-response protocol; 20 runs with reliable outputs form the semantic panel. Pooled Acc@1 is 0.683 and Acc@3 is 0.892, showing that models often locate the correct motivational region while ranking close alternatives unstably. Adjacent values account for 50.9% of semantic errors, compared with 24.4% under a checkpoint-specific null. Eight directed confusions recur across checkpoints and human-confirmed subsets. Several are strongly asymmetric, including Universalism to Benevolence, Tradition to Conformity, and Security to Power, whereas Stimulation-Hedonism forms a bidirectional boundary. Their severity is checkpoint-specific and can bias higher-order value profiles. The results motivate value-recognition evaluation that combines exact accuracy, ranked recovery, and directed error analysis.
发表机构
- Ivannikov Institute for System Programming of the Russian Academy of Sciences(俄罗斯科学院伊万尼科夫系统编程研究所)
- Russian Presidential Academy of National Economy and Public Administration(俄罗斯总统国民经济与公共管理学院)
机构由 AI 辅助整理,请以论文原文为准。