arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Token分布与数据量:多领域会议摘要中的领域平衡

Token Distribution versus Data Volume: Domain Balancing in Multi-Domain Meeting Summarisation

Ashima Sood, Bryan Gardiner, Joan Condell

arXiv 2608.15935首次发表:更新:

发表机构

School of Computing, Engineering and Intelligent Systems, Ulster University(阿尔斯特大学计算、工程与智能系统学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究分离Token分布与数据量因素,在五英文会议语料库上微调Mistral-7B,发现Token平衡可改善少数领域摘要质量,还验证了修剪低价值转写行的作用,为多领域会议摘要的领域平衡提供决策依据。

AI 中文摘要

在规模差异悬殊的会议摘要语料库上对大型语言模型(LLM)进行联合微调时,会产生一个先前研究未厘清的问题:当领域平衡的训练混合数据起作用时,其收益是源于各领域间的Token分布,还是仅仅源于所接触的数据量?我们通过在五个英文会议语料库上构建匹配Token预算(2-32M)的平衡混合数据与自然(原生比例)混合数据来分离这些因素,采用QLoRA对Mistral-7B进行微调,并按领域评估效果。平衡会重新分配质量,在对数据丰富领域造成低损失的同时改善数据稀缺的少数领域。只要少数领域重要,平衡的权衡就有利:无论预算多少,少数领域在原生比例分配中的占比固定为1-2%,因此在这些领域匹配平衡质量需要多得多的总数据。我们还发现,修剪低价值的转写行可从对话语料库中移除约15%的Token且无明显损失,且按Token平衡与按示例平衡并不相同。对741个经标注员标注的事实的双标注员研究验证了我们的事实级评估。这些结果为从业者提供了基础,使其能够决定何时平衡不平衡的多领域混合数据,以及以何种单位进行平衡。

英文摘要

Jointly fine-tuning an LLM on meeting-summarisation corpora of widely varying size raises a question that prior work leaves confounded: when a domain-balanced training mixture helps, is the gain due to the distribution of tokens across domains, or merely to the volume of data seen? We disentangle these factors by constructing balanced and natural (native-proportional) token mixtures at matched token budgets (2-32M) over five English meeting corpora, fine-tuning Mistral-7B with QLoRA, and evaluating per domain. Balancing redistributes quality, improving the data-scarce minority domains at a low cost to the data-rich ones. The trade favours balancing whenever the minority domains matter: their share under proportional allocation is fixed at 1-2% regardless of budget, so matching balanced quality on those domains requires far more total data. We further find that pruning low-value transcript lines removes ~15% of tokens from the conversational corpora at no measurable cost, and that balancing by tokens is not the same as balancing by examples. Fine-tuning one model per domain is competitive only on the data-rich domains and falls below the zero-shot model on the data-scarce ones. A two-annotator study of 741 judge-labelled facts validates our fact-level evaluation. Together these results give practitioners a basis for deciding when to balance an imbalanced multi-domain mixture, and on what unit.

CommentsAccepted at 19th International Natural Language Generation Conference (INLG 2026), Utrecht, Netherlands (camera ready)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑