arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09053cs.AI

身份之后的正义:大语言模型与“无处不在的视角”

Justice After Identity: Large Language Models and the View from Everywhere

W. Russell Neuman

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨大语言模型能否通过编码多样身份视角,近似罗尔斯“原初状态”以实现公平判断,并比较人类与模型在道德困境中的判断差异。

中文摘要 AI 辅助

对正义与公平的共同视角的寻求一直挑战着人类的集体活动,因为我们分歧的判断不可避免地受到社会地位、个人利益、文化传承和历史环境的自我利益所塑造。约翰·罗尔斯曾著名地尝试通过推广一种以“原初状态”为名的哲学传统来克服这一局限——这是一种思想实验,人们在不了解自己将拥有的身份或优势的情况下选择正义原则。然而,批评者长期以来质疑人们是否能够有意义地悬置其社会身份并压制道德相关的生活经验形式。被用于计算算法和智能体公平性的人工智能引入了一种新的可能性。大语言模型没有单一的阶级、种族、性别、国籍或传记,但它们的参数编码了广泛的人类身份和道德传统的语言表征。也许大语言模型的伦理判断可以近似一种整合性的原初状态——一种“无处不在的视角”,它不是通过排除社会身份产生的,而是通过计算性地纳入这些身份产生的。人类不太可能通过简单地将分配性和程序性集体过程的完全智能体控制权委托给计算系统来“交出钥匙”。但人工智能可能发挥作用,也许是积极的作用,与个人和集体的判断互动,因为我们经常面临关于什么公平和正义的日益两极分化的观点。我们提供了数据,比较了人类、基础模型和前沿/微调模型对经典道德困境的判断,同时系统性地变化身份关系和罗尔斯式的身份约束。最后,我们推测,如果先进的人工智能系统为人类提供深思熟虑的建议,人类是否真的可能接受这些建议。

英文摘要

The search for a common view of justice and fairness has challenged human collective activity, as our diverging judgments are unavoidably shaped by the self-interests of social position, personal benefit, cultural inheritance, and historical circumstance. John Rawls famously attempted to overcome this limitation through popularizing a philosophical tradition known by the phrase "the original position" - a thought experiment by which people select principles of justice without knowing the identities or advantages they will possess. Critics, however, have long questioned whether people can meaningfully suspend their social identities and suppress morally relevant forms of lived experience. Artificial intelligence engaged to calculate algorithmic and agentic fairness introduces a novel possibility. LLMs have no singular class, race, gender, nationality, or biography, yet their parameters encode linguistic representations of a vast range of human identities and moral traditions. Perhaps the ethical judgments of LLMs could approximate an integrative original position - a "view from everywhere" generated not by excluding social identities but by computationally incorporating their diversity.It is unlikely that humankind will "hand over the keys" to computational systems by simply delegating complete agentic control of distributive and procedural collective processes. But AI may play a role, perhaps a positive one, interacting with individual and collective human judgment as we often confront increasingly polarized views on what is fair and just. We present data comparing human, base model, and frontier/fine-tuned model judgments about classic moral dilemmas while systematically varying identity relationships and Rawlsian constraints on identity. We conclude by speculating whether, if advanced AI systems provide humans with thoughtful advice, humans would actually be likely to accept it.

↑