通过序列级蒸馏提升小型大语言模型在长法律意见摘要中的论点显著性覆盖率
Improving Argument Saliency Coverage in Small LLMs for Long Legal Opinion Summarization via Sequence-Level Distillation
查看机构详情
- University of Pittsburgh(匹兹堡大学)
- University of the Basque Country (UPV/EHU)(巴斯克大学(UPV/EHU))
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对小型LLMs在长法律意见摘要中难以保留关键论点的问题,提出从高性能长上下文教师模型进行序列级蒸馏的方法,其数据效率高且效果优于专家摘要微调,为该任务提供了高效解决方案。
中文摘要 AI 辅助
我们表明,从高性能长上下文教师模型进行序列级蒸馏是一种简单、无标注且数据高效的策略,可用于提升长法律意见摘要中的论点显著性覆盖率,而小型大语言模型(LLMs)往往难以保留最具显著性的论点内容。在我们的法律意见场景中,针对不同规模的学生模型,蒸馏始终优于基于专家撰写摘要的微调。我们进一步证明,仅需约10条训练摘要即可实现大部分性能提升,凸显了教师生成监督的强数据效率。最后,我们发现摘要蒸馏足以带来性能提升:推理链蒸馏与仅摘要蒸馏表现相当,但结合摘要监督时仅能提供边际收益。
英文摘要
We show that sequence-level distillation from a capable long-context teacher model is a simple, annotation-free, and data-efficient strategy for improving argument saliency coverage in long legal opinion summarization, where small LLMs often struggle to retain the most salient argumentative content. Across student model sizes, distillation consistently surpasses tuning on expert-written summaries in our legal-opinion setting. We further demonstrate that most gains are achieved with as few as ~10 training summaries, highlighting the strong data efficiency of teacher-generated supervision. Finally, we find that summary distillation is sufficient for improvements: reasoning-chain distillation remains competitive with summary-only distillation, but provides marginal benefit when combined with summary supervision.