小模型会遵循你赋予它们的法律吗?孟加拉国法律问答的上下文注入微调
Do Small Models Use the Law You Give Them? Context-Injected Fine-Tuning for Legal QA in Bangladesh
浏览论文内容
中文总结 AI 辅助
研究孟加拉国法律问答,通过整理双语问答记录对不同参数的Qwen3.5进行微调,发现微调能提升部分模型表现,降低答案语言偏差,表明检索质量非唯一瓶颈,小型双语法律模型运用法律和语言回答有差异。
中文摘要 AI 辅助
小语言模型即便接收了适用的法律条文,回答仍可能出错。我们测试了在包含相关法律的示例上进行微调,是否能改善后续对检索到的法律的运用。我们从孟加拉国的六部法案和三个附表中精心整理了2165条双语问答记录,然后对参数为0.8B、2B和4B的Qwen3.5进行微调。评估采用2022年和2023年孟加拉国律师协会考试的孟加拉语及机器翻译英语版本,不使用检索、BM25或FAISS,通过三次种子运行的严格一致性进行评分。在0.8B参数时,微调将2022年英语FAISS分数从100分中的2分提高到34分。0.8B和2B参数模型的提升在配对测试中得以保留,但4B参数模型没有可检测到的净提升:孟加拉语成绩提高,而一些英语条件下的成绩出现倒退。微调还将从孟加拉语偏向英语的答案比例从44.0%-53.2%降至0.2%-0.7%,在每个规模下调整后的p值均小于0.001。因此,检索质量并非唯一瓶颈。小型双语法律模型在运用提供的法律方式以及是否用要求的语言回答方面也存在差异。该数据集可通过此https链接公开获取。
英文摘要
A small language model can receive the governing statutory provision and still answer incorrectly. We test whether fine-tuning on examples containing relevant law improves later use of retrieved law. We curate 2{,}165 bilingual QA records from six Bangladeshi acts and three schedules, then fine-tune Qwen3.5 at 0.8B, 2B, and 4B. Evaluation uses the 2022 and 2023 Bangladesh Bar Council exams in Bangla and machine-translated English, with no retrieval, BM25, or FAISS, scored by strict consistency over three seeded runs. At 0.8B, fine-tuning raises the 2022 English FAISS score from 2 to 34 of 100. Gains at 0.8B and 2B survive paired testing, but the 4B model has no detectable net gain: Bangla improves while several English conditions regress. Fine-tuning also reduces answers that drift from Bangla into mostly English from 44.0--53.2\% to 0.2--0.7\%, with adjusted $p<.001$ at every scale. Retrieval quality is therefore not the only bottleneck. Small bilingual legal models also differ in how they use supplied law and whether they answer in the requested language. The dataset is publicly available at https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.
发表机构
- North South University(南北大学)
- BRAC University(BRAC大学)
机构由 AI 辅助整理,请以论文原文为准。