在挪威MFQ-30上引导LLM响应朝向道德基础
Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30
浏览论文内容
中文总结 AI 辅助
本研究对六个开放权重LLM施测挪威MFQ-30,发现人物角色引导可显著拉近模型与人类道德画像的距离,而ActAdd则使画像变平,并揭示了认知幻影现象。
中文摘要 AI 辅助
近期研究将人类心理测量问卷应用于大型语言模型,以引出道德和价值画像,但尚不清楚这些工具是否测量了模型中任何稳定的东西,或者所得画像能否朝向目标人群移动。我们对六个开放权重LLM施测挪威道德基础问卷(MFQ-30),并将其基础画像与N=1,282名挪威受访者的样本进行比较。我们测试了两种引导干预:提示级人物角色引导和激活级ActAdd。一半模型在我们的注意力检查下参与问卷。另一半默认输出平坦或中心趋势的响应,这些响应在平均上看起来接近人类,但不追踪项目内容。一个中立的北欧受访者人物角色(在编写时未使用人类样本的任何分布信息)使参与模型在Mahalanobis $d^2$上接近挪威均值44-77%。在固定中间层的一对ActAdd使基础画像变平,而非引导各个基础。对于至少一个模型,同一人物角色在改变画像的同时也诱导了基线时缺失的参与,这是Peereboom等人(2025)警告的认知幻影的一个具体实例。
英文摘要
Recent work applies human psychometric questionnaires to large language models to elicit moral and value profiles, but it is not clear whether these instruments measure anything stable in models or whether the resulting profiles can be moved toward a target human population. We administer the Norwegian Moral Foundations Questionnaire (MFQ-30) to six open-weight LLMs and compare their foundation profiles to a sample of N = 1,282 Norwegian respondents. We test two steering interventions, prompt-level persona steering and activation-level ActAdd. Half the models engage with the questionnaire under our attention check. The other half default to flat or central-tendency outputs that look near-human on average without tracking item content. A neutral Nordic-respondent persona, written without any distributional information from the human sample, brings the engaging models 44-77% closer to the Norwegian mean in Mahalanobis $d^2$. One-pair ActAdd at a fixed mid-layer flattens the foundation profile rather than steering individual foundations. For at least one model the same persona that shifts the profile also induces engagement that was absent at baseline, a concrete instance of the cognitive phantoms that Peereboom et al. (2025) warn about.