孟加拉国法律语境下的高效LLM蒸馏:一种兼容智能手机的检索增强生成模型
Efficient LLM Distillation for Bangladesh Legal Context: A Smartphone-Compatible Retrieval-Augmented Generation Model
AI总结:
针对孟加拉国成文法获取难题,通过两阶段蒸馏将90亿参数教师模型压缩为20亿参数学生模型,结合混合检索实现智能手机端离线运行,显著提升回答质量并获律师认可。
AI中文摘要:
孟加拉国的法律信息对大多数公民而言难以获取。成文法仅为英文,训练有素的律师集中于城市中心,而依赖云端的AI在移动连接不可靠的环境中失效,在这种环境下,幻觉法律文本会造成直接伤害。该系统仅处理成文法解释;需要司法先例或判例法推理的查询不在其范围内。我们通过两阶段渐进式知识蒸馏,将90亿参数的Gemma-2教师模型压缩为20亿参数的学生模型,以解决成文法获取差距。第一阶段在9,429个质量门控的法律问答对上进行监督微调(从14,514个生成查询中接受率为65%);第二阶段在温度tau=4.0下,最小化与教师模型每词元前50个logits的稀疏Kullback-Leibler散度,通过QLoRA(4位NF4,秩32的LoRA适配器)实现。先前的法律语言模型针对通用法律英语;本系统专精于孟加拉国成文法。每个回答均通过混合检索进行接地,该检索结合了密集语义搜索(60%)和BM25(40%),覆盖来自孟加拉国宪法和国家立法的36,029个成文段落。在50个查询的英文基准上,蒸馏模型达到ROUGE-L 0.4715和BERTScore F1 0.5679,相比检索增强的未蒸馏基线(ROUGE-L 0.2323,BERTScore 0.2340)分别提升103%和143%。适配器量化至1.6 GB(GGUF Q4_K_M),在Pixel 6上以每秒4-8个词元的速度运行,无需网络访问。对50个孟加拉语查询的跨语言评估产生ROUGE-L 0.4083和BERTScore 0.8133,表明从孟加拉语输入针对纯英文语料库的检索有效。在单一评估者试点中,一名执业律师对50个回答的加权平均评分为4.16/5(90%评为4或5),支持超越文本重叠指标的实用性。
英文摘要:
Legal information in Bangladesh is inaccessible to most citizens. Statutory text is English-only, trained lawyers are concentrated in urban centres, and cloud-dependent AI fails where mobile connectivity is unreliable, a setting in which hallucinated legal text causes direct harm. The system addresses statutory interpretation only; queries that require judicial precedent or case-law reasoning fall outside its scope. We target the statutory access gap by compressing a 9-billion-parameter Gemma-2 teacher into a 2-billion-parameter student through two-phase progressive knowledge distillation. Phase 1 performs supervised fine-tuning on 9,429 quality-gated legal question-answer pairs (65% acceptance from 14,514 generated queries); Phase 2 minimises sparse Kullback-Leibler divergence against the teacher's top-50 per-token logits at temperature tau = 4.0, implemented via QLoRA (4-bit NF4, rank-32 LoRA adapters). Prior legal language models target general legal English; this system specialises in Bangladeshi statutory law. Every response is grounded through hybrid retrieval combining dense semantic search (60%) and BM25 (40%) across 36,029 statutory passages from the Bangladesh Constitution and national legislation. On a 50-query English benchmark, the distilled model reaches ROUGE-L 0.4715 and BERTScore F1 0.5679, a 103% ROUGE-L and 143% BERTScore gain over the retrieval-augmented undistilled baseline (ROUGE-L 0.2323, BERTScore 0.2340). The adapter quantises to 1.6 GB (GGUF Q4_K_M) and runs at 4-8 tokens per second on a Pixel 6 with no network access. Cross-lingual evaluation on 50 Bangla queries yields ROUGE-L 0.4083 and BERTScore 0.8133, showing effective retrieval from Bangla input against an English-only corpus. In a single-evaluator pilot, a practising lawyer rated 50 responses at a weighted mean of 4.16/5 (90% rated 4 or 5), supporting utility beyond text-overlap metrics.