arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

引导大型语言模型遵循法律的精神或字面含义

Directing large language models to follow the letter or spirit of the law

Peng Qian, Andrew Li, Sam Chen, Sonia K. Murthy, Yonatan Belinkov, Tomer D. Ullman

arXiv 2609.23083首次发表:更新:

发表机构

Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University; Taub Faculty of Computer Science, Technion — Israel Institute of Technology; Department of Psychology, Harvard University(哈佛大学肯普纳自然与人工智能研究所; 以色列理工学院陶布计算机科学学院; 哈佛大学心理学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过针对性适配引导大型语言模型优先遵循法律精神或字面含义,以最小修改显著改变其行为,并揭示内部低维空间中的三个可解释维度,展示了法律思维的引导机制。

AI 中文摘要

法律精神与字面含义之间的区别是研究和日常生活中的核心问题,也是构建安全、智能机器日益关注的焦点。这种区别基于什么,我们如何开发能够遵循规则背后意图的机器?我们采用了有针对性的适配方法,使大型语言模型优先考虑法律的精神或字面含义。通过最小的修改,我们的方法在多种度量、新颖场景、现实世界情境和有影响力的法律案例中显著改变了大型语言模型的行为。对模型内部的分析揭示了一个低维空间,其中包含三个可解释的维度,与法律概念几何的正式预指定框架相匹配。这些发现展示了大型语言模型中的法律思维如何被组织和引导。

英文摘要

The distinction between the spirit and letter of the law is a central issue across research and everyday life, and a growing concern for building safe, intelligent machines. What is this distinction based on, and how can we develop machines that follow the intention behind a rule? We used targeted adaptation that made large language models prioritize the spirit or letter of the law. With minimal modifications, our method significantly changed LLM behavior across diverse measures, novel vignettes, real-world scenarios, and influential legal cases. An analysis of model internals revealed a low-dimensional space with three interpretable dimensions matching a formal pre-specified framework for the geometry of legal concepts. These findings show how legal thought in LLMs may be organized and directed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑