arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

俄语ModernBERT在长法律文档上的领域适配

Domain adaptation of Russian ModernBERT for long legal documents

I. Litvak, D. Gvozdetsky, F. Lashkin, V. Kirova, S. Lagutin, V. Volf, T. Maksiyan, A. Kostin, R. Leva, I. Kiselev

arXiv 2610.06715首次发表:更新:

发表机构

Independent Researcher; National Research Nuclear University MEPhI; Moscow Center for Advanced Studies; University of Waterloo(独立研究者; 国立核研究大学 MEPhI; 莫斯科高级研究中心; 滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过继续预训练构建俄语法律领域模型RuModernBERT-ruLaw,在长法律文档上降低掩码词元预测损失,但实体抽取评估受数据重叠限制,迁移证据有限。

AI 中文摘要

我们研究了在俄语立法文档上继续预训练是否能提升俄语ModernBERT编码器在法律文本上的表现。适配后的模型RuModernBERT-ruLaw在据报道包含304,382份立法文档和194,425,905个语料库词元的语料库上进行了训练。语料库词元计数与模型分词器产生的位置区分开来。我们在一个固定的外部集合(包含1,031个法院判决片段)上比较了原始编码器和适配后的编码器。两个模型在五种掩码实现中接收相同的隐藏位置。在最大输入长度为512、2,048和8,192个词元时,平均掩码词元交叉熵分别降低了0.10942、0.07052和0.06604自然对数单位。报告的95%区间总结了该固定集合上对掩码的敏感性;它们并不量化跨文档集合的不确定性。第二个评估涉及法律实体抽取。原始模型和适配后的模型分别达到了0.99852和0.99820的实体级F1分数。然而,测试跨度中99.95%在训练分割中具有相同的归一化表面形式和类别。因此,该评估对向未见形式的迁移提供的证据有限。本文使用可编辑的图表和清晰标记的说明性示例解释了掩码目标、重叠窗口、平均规则和精确实体边界评分。该比较支持所研究的模型对和集合上较低的掩码词元预测损失。它并未孤立出远距离上下文的贡献,也未确立实际的法律实用性。

英文摘要

We investigate whether continued pretraining on Russian legislative documents improves a Russian ModernBERT encoder on legal text. The adapted model, RuModernBERT-ruLaw, was trained on a corpus reported to contain 304,382 legislative documents and 194,425,905 corpus tokens. Corpus token counts are distinguished from positions produced by the model tokenizer. We compare the original and adapted encoders on a fixed external collection of 1,031 court-decision segments. Both models receive the same hidden positions in each of five masking realizations. At maximum input lengths of 512, 2,048, and 8,192 tokens, mean masked-token cross-entropy decreases by 0.10942, 0.07052, and 0.06604 natural-log units, respectively. The reported 95% intervals summarize sensitivity to masking on this fixed collection; they do not quantify uncertainty across document collections. A second evaluation addresses legal-entity extraction. The original and adapted models achieve entity-level F1 scores of 0.99852 and 0.99820. However, 99.95% of test spans have the same normalized surface form and class in the training split. This evaluation therefore provides limited evidence about transfer to previously unseen forms. The paper explains the masking objective, overlapping windows, averaging rules, and exact entity-boundary scoring using editable diagrams and clearly marked illustrative examples. The comparison supports lower masked-token prediction loss for the studied pair of models and collection. It does not isolate the contribution of distant context or establish practical legal utility.

Comments17 pages, 11 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑