AI 中文总结
本研究通过意大利律师、法官、公证人考试的盲法图灵测试,评估开箱即用的主流LLM的法律能力,发现其在部分法律任务表现接近人类但公证人考试全部失败,明确了LLM法律能力的范围与边界。
AI 中文摘要
本文报告了一项盲法图灵测试实验,评估开箱即用的主流大语言模型(LLM)在三项意大利法律职业考试(律师考试、法官考试、公证人考试)中的表现。要求主流LLM生成完整的书面考试答卷,这些答卷需与人类提交的答卷难以区分,并由专业考官采用与实际考试相同的标准进行匿名评估。结果显示,不同模型及不同任务间存在显著差异:部分LLM在对抗性法律论证和教义分析中达到或超过顶尖人类表现,所有模型均在公证人考试中失败,该考试要求在严格的形式和实质约束下进行目标导向的法律规划。除对模型进行排名外,本研究还确定了特定任务的优势、局限性及反复出现的法律失败模式。尽管研究仅针对开箱即用的系统,其发现仍为LLM在不同专业任务中法律能力的当前范围和边界提供了定性证据。
英文摘要
The article reports on a blind Turing Test experiment, assessing the performance of out-of-the-box leading LLMs on three Italian legal professional exams: the Bar, Judges and Notary exams. Leading LLMs were asked to generate full written exam papers, which were made indistinguishable from human submissions and anonymously evaluated by expert examiners, using the same criteria applied in real examinations. Results reveal marked differences across both models and tasks. While some LLMs match or exceed top human performance in adversarial legal argumentation and doctrinal analysis, all models fail in the notary exam, which requires goal-directed legal planning under strict formal and substantive constraints. Beyond ranking models, the study identifies task-specific strengths, limitations and recurring legal failure patterns. Although limited to out-of-the-box systems, the findings provide qualitative evidence on the current scope and boundaries of the legal competence of LLMs across distinct professional tasks.