发表机构
Rutgers University; Pennsylvania State University; Binghamton University(罗格斯大学; 宾夕法尼亚州立大学; 宾汉姆顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种通用方法评估大型语言模型在资源分配中的公平推理,发现其比人类更偏好严格约束、更自利且难以通过现有数据集微调对齐人类判断。
AI 中文摘要
稀缺且不可分割资源的公平分配是许多社会问题中的重要挑战。尽管存在多种正式的公平理论,但没有单一定义总能被满足。随着大型语言模型(LLMs)日益被用于支持决策并充当智能体,它们引发了关于分配正义的新担忧:其判断并不直接关联于任何特定的公平框架,且可能违反关键的规范性原则。在本工作中,我们引入了一种评估LLMs公平推理的通用方法。我们研究了一系列模型中的第一人称公平判断,并将其与人类在匹配场景和诱导条件下的反应直接比较。我们发现,LLMs倾向于偏好比人类更严格的公平约束,表现出更自利的行为,对信息框架方式敏感,并且难以通过使用当前数据集的微调来与人类判断对齐。
英文摘要
Fair allocation of scarce, indivisible resources is an important challenge in many societal problems. While there are several formal theories of fairness, no single definition can always be satisfied. As large language models (LLMs) are increasingly used to support decisions and act as agents, they raise new concerns about distributional justice: their judgments are not directly tied to any specific fairness framework and may violate key normative principles. In this work, we introduce a general method for evaluating fairness reasoning in LLMs. We study first-person fairness judgments across a broad set of models and compare them directly with human responses on matched scenarios and elicitation conditions. We find that LLMs tend to prefer stricter fairness constraints than humans, show more self-interested behavior, are sensitive to how information is framed, and are difficult to align with human judgments using fine-tuning with current datasets.
CommentsAccepted at EMNLP 2026 (Main Conference)