arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26945cs.CLcs.CY

句子层面的解释准则分类:来自德国联邦宪法法院的基准

Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court

Felix Ringe

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建句子级基准,评估大语言模型分类德国宪法法院解释准则的能力,发现文义解释易、体系解释难,专家提示优于GEPA优化提示。

中文摘要 AI 辅助

司法推理仍然是大语言模型(LLMs)难以分析的挑战。本文贡献了一个句子层面的基准,用于评估LLMs对Larenz在Savigny传统中阐述的解释准则进行分类的能力。我们的贡献有三方面。首先,我们将这种解释概念操作化为分类标准。其次,我们提供了一个在句子层面标注的德国联邦宪法法院判决数据集。第三,我们报告了来自三个模型家族的四个LLM在专家手写提示下的基线评估,并与使用遗传-帕累托(GEPA)优化的提示进行比较。七个二元子任务的平均F1分数在模型间聚集于70.4至79.2之间,其中文义解释通常是最容易识别的准则,而体系解释通常是最难的;在测试配置下,GEPA优化的提示并未系统性地优于手写提示,这表明专家提示提供了一个有意义的基线。

英文摘要

Judicial reasoning remains challenging for large language models (LLMs) to analyze. This paper contributes a sentence-level benchmark for evaluating the ability of LLMs to classify interpretive canons as articulated by Larenz in the tradition of Savigny. Our contributions are threefold. First, we operationalize this conception of interpretation as classification criteria. Second, we provide a dataset of decisions of the German Federal Constitutional Court annotated at the sentence level. Third, we report baseline evaluations of four LLMs from three model families under expert hand-written prompts, compared against prompts optimized with Genetic-Pareto (GEPA). Mean F1 over the seven binary subtasks clusters between 70.4 and 79.2 across models, with grammatical interpretation usually the easiest canon to identify and systematic interpretation usually the hardest; under the tested configuration, GEPA-optimized prompts do not systematically outperform the hand-written ones, suggesting that the expert prompts provide a meaningful baseline.

发表机构

  • Freie Universität Berlin(柏林自由大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑