arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MIDAS:多大语言模型迭代数据自适应摘要生成

MIDAS: Multi-LLM Iterative Data-Adaptive Summarization

Karen Lee, Dhanashree Balaram, Seojun Shon, Umair Rasheed

arXiv 2608.04307首次发表:更新:

发表机构

Volkswagen Group Innovation(大众集团创新公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有自动提示词优化方法无法适配多样摘要需求的问题,提出多大语言模型迭代数据自适应摘要生成框架MIDAS,在企业客户工单摘要任务中优于CriSPO等方法,还具备跨模型与跨领域泛化能力。

AI 中文摘要

文本摘要的难度超乎想象。虽然浓缩信息看似简单,但现实中企业对支持工单、法律文件、事件报告等内容的摘要,要求严格遵循领域特定准则、输出格式和组织惯例。制作能可靠满足这些约束的提示词十分费力,需要丰富的专业知识,且随着需求演变还要持续维护。现有的自动提示词优化方法通过大语言模型(LLM)批评驱动的优化减轻了这一负担,但仍受限于静态提示词,无法适应不同摘要应用的多样性。我们提出多大语言模型迭代数据自适应摘要生成框架(MIDAS),该多大语言模型框架将这一范式扩展为数据驱动的模式学习和用例特定个性化,无需手动提示词工程即可自动适配不同的摘要需求。将其应用于五种输出格式的企业客户工单摘要任务时,MIDAS在CriSPO、ZERA等最先进的批评驱动优化框架中实现了最强的整体性能,ROUGE-1最高提升11.0%,ROUGE-2最高提升18.2%,ROUGE-L最高提升8.0%,且在所有格式和输出类型中均持续提升BERTScore F1。我们还通过多大语言模型配置和金融领域摘要基准验证了其跨模型和跨领域泛化能力。

英文摘要

Text summarization is deceptively difficult. While condensing information seems straightforward, real-world enterprise summarization of support tickets, legal documents, incident reports, and more, demands strict adherence to domain-specific guidelines, output formats, and organizational conventions. Crafting prompts that reliably satisfy these constraints is labor-intensive, requiring significant human expertise and continuous maintenance as requirements evolve. Existing automated prompt optimization methods reduce this burden through Large Language Model (LLM) critique-driven refinement, yet remain limited by static prompts that cannot adapt to the diversity of summary applications. We propose Multi-LLM Iterative Data-Adaptive Summarization (MIDAS), a multi-LLM framework that extends this paradigm with data-driven pattern learning and use-case-specific personalization, enabling automatic adaptation to different summarization requirements without manual prompt engineering. Applied to enterprise customer ticket summarization across five output formats, MIDAS achieves the strongest overall performance against state-of-the-art critique-driven optimization frameworks such as CriSPO and ZERA, improving ROUGE-1 by up to 11.0%, ROUGE-2 by up to 18.2%, and ROUGE-L by up to 8.0%, while consistently improving BERTScore F1 across all formats and output types. We additionally demonstrate cross-model and cross-domain generalization through multi-LLM configurations and finance-domain summarization benchmarks.

CommentsAccepted at the 20th International Conference on Document Analysis and Recognition (ICDAR 2026). 17 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑