发表机构
North South University; BRAC University; HuggingFace(北南大学; BRAC大学; HuggingFace)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对双语孟加拉法律问答任务,发现微调未提升小模型对上下文中法律的依赖,需分离评分、检索及模型效应对评估法律适配性的影响
AI 中文摘要
微调可提升法律问答的准确率,但不会提升模型对上下文中提供的法律的利用程度。我们在孟加拉语-英语双语孟加拉法律问答任务中研究了这一区别,观察到的错误可能源于答案评分、检索或未能利用相关法律。我们构建了保留层级结构的法规语料库、2165个经审查的双语微调示例,以及150项提供法律的对照项。我们评估了6个指令微调模型:Llama-3.2-1B、Llama-3.2-3B、Qwen3.5-0.8B、Qwen3.5-2B、Qwen3.5-4B和Gemma-4-E2B,每个模型有3个LoRA种子。为分离效应,我们结合了受限选项字母评分、循环选项旋转以及对管辖条款的受控移除。在398个Bar Council输出上,精确行解析器将Qwen3.5-2B seed-42适配器的准确率提升归因于50.0%,而选项评分仅产生3.0%的提升;对于Gemma-4-E2B,两种评分方法倾向于不同的系统。当管辖条款被确保存在时,6个参考模型中有5个在四阶准则下提升了14.7%-19.3%;移除该条款会使模型准确率降低8.0%-15.3%,其适配器准确率降低13.8-14.9个百分点。然而,双重差分估计显示,微调后模型对管辖条款的依赖并未增加。结果表明,法律适配性主张需要分离评分器、检索器和模型的效应。我们的代码和数据可在该https URL获取
英文摘要
Fine-tuning can improve legal question-answering accuracy without improving how models use law supplied in context. We study this distinction in bilingual Bangladeshi legal QA, where observed errors can arise from answer scoring, retrieval, or failure to use relevant law. We construct a hierarchy-preserving statutory corpus, 2,165 reviewed bilingual fine-tuning examples, and a 150-item supplied-law control. We evaluate six instruction-tuned models: Llama-3.2-1B, Llama-3.2-3B, Qwen3.5-0.8B, Qwen3.5-2B, Qwen3.5-4B, and Gemma-4-E2B, with three LoRA seeds per model. To separate effects, we combine constrained option-letter scoring, cyclic option rotation, and controlled removal of the governing provision. On 398 Bar Council outputs, an exact-line parser attributes an accuracy gain of 50.0\% to the Qwen3.5-2B seed-42 adapter, whereas option scoring yields only $3.0\%$. For Gemma-4-E2B, the two scoring methods favor different systems. When the governing provision is guaranteed to be present, five of six reference models improve by $14.7\%-19.3\%$ under the four-order criterion. Removing that provision reduces accuracy by $8.0\%-15.3\%$ for models and by $13.8\%-14.9\%$ points for their adapters. However, difference-in differences estimates show no increase in reliance on the governing provision after fine-tuning. Results show that legal adaptation claims require separating scorer, retriever, and model effects. Our Code and data are available at https://anonymous.4open.science/r/bangladesh-legal-qa-11E3
CommentsLegal Data Benchmark for Bangladesh