arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06769cs.CYcs.AI

普通、合理的聊天机器人:AI模型是否追踪人类法律判断?

Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?

  • Duke University(杜克大学)
  • Duke University Law School(杜克大学法学院)

机构由 AI 辅助整理,请以论文原文为准。

Nirav Patel, Emily Wenger, Christopher Buccafusco

AI总结:

本研究比较26个LLM与人类在25项法律合理性判断上的回答,发现聊天机器人总体追踪人类判断,但更同质化、偏向政府和公司,且与特定人口群体更一致。

AI中文摘要:

随着人们越来越依赖人工智能(AI)来指导自己的生活,学者、律师甚至法官开始考虑AI在法律决策中的作用。正如“硅基抽样”——即在社会科学研究中使用生成式AI模型——如今正影响着学术界,“硅基陪审员”也可能出现在法庭上。本研究加入了关于生成式AI模型模拟人类法律判断能力的新兴研究行列。具体而言,我们研究了由大语言模型(LLM)驱动的聊天机器人如何回应一系列关于法律合理性的问题。当法律需要判断行为的适当性时,它最常问的是该行为是否“合理”。然而,尽管合理性判断无处不在,它们却是律师、法官和外行人不断感到困扰的所在。合理性似乎天生模糊且不可预测,因为它依赖于多变的背景和隐含的概念图式。此外,许多学者警告说,合理性判断可能因人口统计学特征而异。我们比较了人类参与者与二十六个LLM在二十五项不同的法律相关合理性判断上的回答。总体而言,我们的研究结果表明,聊天机器人的回答通常与人类参与者的回答一致。尽管如此,我们确实发现了一些具有暗示性——且可能令人担忧——的结果。与人类相比,LLM产生了更加同质化的回答,并且偶尔会将可变的标准视为不变的规则。而且,与人类相比,LLM倾向于生成对政府和公司更有利的回答。最后,我们的结果表明,LLM的回答往往与那些白人、男性、年龄较大且受教育程度较高的受访者的回答更为接近。需要更系统的研究来确认或否定这些初步发现。

英文摘要:

As people increasingly rely on artificial intelligence (AI) for guidance in their own lives, scholars, lawyers, and even judges have begun to consider the role of AI in legal decision-making. As "silicon sampling" -- the use of generative AI models in social science research -- is now impacting academia, "silicon jurors" could make an appearance in courtrooms. This study joins an emerging line of research on generative AI models' ability to simulate human legal judgments. In particular, we study how large language model (LLM)-powered chatbots respond to series of questions about legal reasonableness. When the law needs to judge the appropriateness of a behavior, it most often asks whether the behavior was "reasonable." Yet despite the ubiquity of reasonableness judgments, they are the site of constant vexation for lawyers, judges, and lay people. Reasonableness seems inherently vague and unpredictable, since it relies on variable context and implicit conceptual schemas. Moreover, many scholars caution that reasonableness judgments may vary along demographic lines. We compare the answers of human participants to those of twenty-six LLMs across twenty-five different legally relevant reasonableness judgments. Overall, our findings suggest that chatbot responses generally track those of human participants. Nonetheless, we find some suggestive -- and potentially concerning -- results. Compared to humans, LLMs generate more homogeneous responses and occasionally treat a variable standard as an invariant rule. And, compared to humans, LLMs tend to generate answers that are more favorable to the government and to corporations. Finally, our results indicate that LLMs' responses tend to align more closely with those of respondents who are white, male, older, and more educated. More systematic research is needed to confirm or reject these initial findings.

补充信息

↑