AI 中文总结
本文针对德国法定文本逻辑成分提取难题,构建ANNOTARES数据集并测试多种模型,发现BERT与LLM模型表现更优,将发布数据集以推动自动法律推理研究。
AI 中文摘要
法律文本的自动结构分析是法律技术的基石,但其逻辑成分的提取仍是一项重大挑战。本文提出了识别和分割德国法定文本中法律要件(Tatbestand)与法律后果(Rechtsfolge)的任务。为支持该任务,我们推出ANNOTARES(即法律要件-法律后果序列标注数据集),这是一个包含德语法律文本及跨度级标注的新型数据集,涵盖三部不同的法典,旨在评估领域特定性能及跨法典泛化能力。我们对多种架构方法进行基准测试:基于规则的基线、条件随机场(CRFs)、双向长短期记忆网络(BiLSTMs)、BiLSTM-CRF,以及现代基于Transformer的模型,包括BERT变体和基于大语言模型(LLM)的方法。结果表明,BERT和基于LLM的模型在捕捉法律语言复杂句法结构方面表现更优。我们发布该数据集以推动自动法律推理领域的进一步研究。
英文摘要
The automatic structural analysis of legal texts is a cornerstone of legal technology, yet the extraction of their logical components remains a significant challenge. In this paper, we introduce the task of identifying and segmenting legal conditions (Tatbestand) and legal consequences (Rechtsfolge) within German statutory texts. To support this task, we present ANNOTARES (Annotations of Tatbestand-Rechtsfolge Sequences), a novel dataset comprising German law texts with span-level annotations. Spanning three distinct legal codes, the dataset is designed to evaluate both domain-specific performance and cross-statute generalizability. We benchmark diverse architectural approaches: a rule-based baseline, CRFs, BiLSTMs, BiLSTM-CRF, and modern Transformer-based models, including BERT variants and LLM-based methods. Our results demonstrate that BERT and LLM-based models achieve superior performance in capturing the complex syntactic structures of legal language. We release our dataset to facilitate further research in automated legal reasoning.
CommentsAccepted at KONVENS 2026