发表机构
JustAI; INSA Rouen Normandie(JustAI; 鲁昂诺曼底国立应用科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有AI对齐方法的局限,提出法定人工智能,以法律文本为框架,通过两阶段思维链提示运行,在1000个红队提示的五主题实验中,较标准宪法AI更能减少有害内容并缩短计算时间。
AI 中文摘要
随着AI监管框架的不断发展,确保人工智能系统,尤其是生成式模型,按照法律和伦理标准运行已成为关键优先事项。然而,现有的AI对齐和价值引导行为方案存在一些局限性。诸如宪法AI(Constitutional AI)这类方法依赖人工监督,而像“有益于人类”(Good-for-Humanity, GfH)原则这类宽泛的规范框架可能过于笼统和模糊,无法提供可操作的治理指导。为克服这些局限,我们提出一种名为法定人工智能(Statutory AI)的混合方法,该方法采用从法律语料库特定主题中提取的现有人类撰写的原则。具体而言,法定人工智能将法律文本用作宪法框架,使AI系统能够依据既定规范自主批判和修订自身输出。该方法分两个阶段运行,两个阶段均使用思维链(Chain-of-Thought)提示。第一阶段将用户提示分类到已识别的主题之一,第二阶段结合该主题法律语料库中选定的相关条款对提示进行分析。为说明我们方法的潜力,我们开展了一项涉及1000个红队提示和五个惩罚性主题的实验,这五个主题分别是歧视、机密信息泄露、暴力、欺诈以及对弱势群体的虐待。法定人工智能在测试模型上将有害内容减少了52至59个百分点,比标准宪法AI(Constitutional AI)高出约10个百分点,同时将计算时间缩短了50%以上。
英文摘要
With the increasing development of AI regulatory frameworks, ensuring that artificial intelligence systems, particularly generative models, operate in accordance with legal and ethical standards has become a critical priority. Existing proposals for AI alignment and value-guided behavior, however, face some limitations. Approaches such as Constitutional AI depend on human supervision, while broad normative frameworks like the Good-for-Humanity (GfH) principle may be overly general and ambiguous to provide actionable governance guidance. To overcome these limitations, we propose a hybrid approach called Statutory AI that employs pre-existing human-authored principles drawn from specific themes within a legal corpus. Specifically, Statutory AI uses legal texts as a constitutional framework, enabling AI systems to autonomously critique and revise their outputs according to established norms. It operates in two stages, both using Chain-of-Thought prompting. The first stage classifies the user prompt into one of the identified themes, while the second stage analyzes it in conjunction with relevant articles selected from the legal corpus of that theme. To illustrate the potential of our approach, we conducted an experiment involving 1,000 red-teaming prompts and five penal themes: discrimination, disclosure of confidential information, violence, fraud, and abuse of vulnerable persons. Statutory AI reduced harmful content by 52 to 59 percentage points across tested models, approximately 10 percentage points higher than standard Constitutional AI, while cutting computation time by over 50%.
Journal refArtificial Intelligence for Digital Transformations (AIDT 2026), Communications in Computer and Information Science, vol. 3023, pp. 375-388, Springer, Cham, 2026
DOI:10.1007/978-3-032-31319-5_25