arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GreenLeaf Law Embed Tiny:一款面向法律领域检索的紧凑嵌入模型

GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval

Surya Saka

arXiv 2608.24936首次发表:更新:

发表机构

JudicialMind(司法思维机构)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出参数0.6B的GreenLeaf Law Embed Tiny法律领域检索嵌入模型,通过两阶段训练等方法,在相关基准取得优异成绩,可部署于资源受限环境。

AI 中文摘要

我们提出GreenLeaf Law Embed Tiny,这是一款拥有0.6B参数的法律领域检索嵌入模型。GreenLeaf-Tiny在大规模法律嵌入基准(Massive Legal Embedding Benchmark, MLEB)上取得75.11%的成绩,在MTEB(Law, v1)上取得64.38%的成绩,在参数小于1B的模型中展现出有竞争力的性能。我们的方法结合了两阶段训练流程:首先将较大的教师模型的知识蒸馏到紧凑的学生架构中,随后利用难负例挖掘进行领域特定微调;我们还使用了精心整理的数据集,包含340万条查询-段落对,其中涵盖15万条来自不同司法管辖区的人工整理样本;此外,我们采用了支持多种量化级别(BF16、INT8、二进制)的高效推理架构,可在资源受限的环境中部署。我们对训练方法、架构选择以及法律检索任务上的综合评估进行了详细分析,结果表明,利用高质量数据进行领域特定训练可提升专业领域应用的性能。

英文摘要

We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval. GreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1),demonstrating competitive performance among models under 1B parameters. Our approach combines a two-stage training pipeline that first distills knowledge from a larger teacher model into a compact student architecture, then applies domain-specific fine-tuning with hard negative mining; a carefully curated dataset of 3.4 million query-passage pairs, including 150,000 human-curated samples across diverse legal jurisdictions; and an efficient inference architecture supporting multiple quantization levels (BF16, INT8, binary) enabling deployment in resource-constrained environments. We provide detailed analysis of our training methodology, architectural choices, and comprehensive evaluation across legal retrieval tasks. Our results demonstrate that domain-specific training with high-quality data can improve performance for specialized domain applications

Comments7 pages, 2 figures, IEEE dual-column format. Submitted to arXiv for preprint distribution

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑