SinBrief:一种用于僧伽罗语法律文档的混合式抽象文本摘要框架
SinBrief: A Hybrid Framework for Abstractive Text Summarisation of Sinhala Legal Documents
AI总结:
SinBrief提出一种无需人工标注的混合抽象摘要框架,结合词图构建与神经句子评分,在僧伽罗语法律文本上实现低词汇重叠且保持事实一致性的摘要生成。
AI中文摘要:
低资源语言中的法律文档摘要因标注数据稀缺和领域特定术语复杂而面临重大挑战。本文提出了SinBrief,一种针对僧伽罗语法律文档的混合式抽象摘要框架,该框架无需人工标注的训练数据。所提出的框架将领域感知的词图构建与神经句子评分相结合,从僧伽罗语法律文本中生成抽象式摘要。框架内评估了五种句子评分模型:mBert、Llama 3.1、Falcon 7B、Laser以及一个持续预训练并针对僧伽罗语法律文本进行领域适配的Llama模型。该框架在僧伽罗语法律语料库上使用无参考指标进行评估,包括Coverage、Density、Compression Ratio、SummaC和Self-BertScore。实验结果表明,SinBrief生成的摘要与抽取式基线相比具有更低的词汇重叠度,同时保持事实一致性,证明了混合式、基本免标注的抽象摘要方法在低资源法律NLP任务中的可行性。
英文摘要:
Legal document summarisation in low-resource languages presents significant challenges due to the scarcity of annotated data and the complexity of domain-specific terminology. This paper presents SinBrief, a hybrid abstractive summarisation framework for Sinhala legal documents that does not require human-annotated training data. The proposed framework combines domain-aware word graph construction with neural sentence scoring to generate abstractive summaries from Sinhala legal text. Five sentence scoring models are evaluated within the framework: mBert, Llama 3.1, Falcon 7B, Laser, and a continually pre-trained Llama model domain-adapted to Sinhala legal text. The framework is evaluated on a Sinhala legal corpus using reference-free metrics, including Coverage, Density, Compression Ratio, SummaC, and Self-BertScore. Experimental results demonstrate that SinBrief produces summaries with lower lexical overlap than extractive baselines while maintaining factual consistency, demonstrating the viability of hybrid, largely annotation-free abstractive summarisation for low-resource legal NLP tasks.