arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

韩国性犯罪案件中的法律文本分类:从传统机器学习到具有XAI洞察的大型语言模型

Legal text classification in Korean sexual offense cases: from traditional machine learning to large language models with XAI insights

Jeongmin Lee

arXiv 2610.00087首次发表:更新:

AI 中文总结

本研究评估韩国性犯罪法律文本分类模型,发现微调KLUE-BERT达99.3%准确率,优于GPT-3.5等大型模型,并利用XAI分析揭示模型决策特征与局限,强调可解释性在法律AI中的重要性。

AI 中文摘要

自然语言处理(NLP)的进步扩展了法律领域中基于AI的文本分类应用。然而,由于法律文本的复杂性和法律类别之间的细微差异,准确分类法律文件仍然具有挑战性。本研究使用韩国性犯罪先例的十个类别,评估了从传统机器学习技术到大型语言模型(LLMs)的法律文本分类模型。结果表明,在法律数据上微调KLUE-BERT等小规模模型,其性能优于GPT-3.5和GPT-4.0等通用模型以及传统机器学习模型。KLUE-BERT达到了99.3%的最高准确率,表明对于法律文件分类,领域适应和微调可能比模型规模更为重要。我们进一步采用可解释人工智能(XAI)技术来分析模型预测和错误分类案例。XAI分析识别出影响模型决策的语言特征,以及在捕捉细微文本线索方面的局限性。使用与真实世界法律案件记录高度相似的KICS数据,我们进一步评估了模型的泛化能力,并发现其在解释隐含上下文线索方面存在困难。这些发现强调了在法律AI中性能和可解释性的重要性,并展示了XAI如何提高法律文本分类的透明度。AI辅助工具可以支持法律专业人员完成文件分类、法律信息检索和案件评估等任务。

英文摘要

The advancement of natural language processing (NLP) has expanded AI-based text classification in the legal domain. However, accurately classifying legal documents remains challenging due to the complexity of legal texts and subtle differences between legal categories. This study evaluates legal text classification models ranging from traditional machine learning techniques to large language models (LLMs) using ten categories of Korean sexual offense precedents. The results show that fine-tuning small-scale models such as KLUE-BERT on legal data outperforms general-purpose models such as GPT-3.5 and GPT-4.0, as well as traditional machine learning models. KLUE-BERT achieved the highest accuracy of 99.3%, indicating that domain adaptation and fine-tuning can be more important than model size for legal document classification. We further employ explainable AI (XAI) techniques to analyze model predictions and misclassification cases. XAI analysis identifies linguistic features influencing model decisions and limitations in capturing subtle textual cues. Using KICS data, which closely resembles real-world legal case records, we further evaluate the model's generalization capabilities and find that it struggles to interpret implicit contextual cues. These findings highlight the importance of both performance and interpretability in legal AI and demonstrate how XAI can improve transparency in legal text classification. AI-assisted tools can support legal professionals in tasks including document classification, legal information retrieval, and case assessment.

Comments22 pages, 3 figures. Published in Artificial Intelligence and Law

Journal refJ. Lee, Artificial Intelligence and Law (2025)

DOI:10.1007/s10506-025-09454-w

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑