arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MMLDSum-LLM:结合视觉对齐与关键词感知的多模态长文档摘要方法

MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware

Xianpeng Zhang, Jiahua Yang, Dongyu Chen, Lei zhang, Jian Ma, Xu guohuan, Haonan Lu, Tianhuang Su, Chuangchuang Wang, Kai Tang

arXiv 2607.28006首次发表:更新:

AI 中文总结

针对多模态长文档摘要的关键信息遗漏与跨模态幻觉问题,提出结合视觉对齐与关键词感知的两阶段训练框架MMLDSum-LLM,在自研基准MMLDSum-Bench上验证了其性能优势。

AI 中文摘要

多模态长文档是专业知识的核心载体,关键证据稀疏分布在段落与模态中,易导致多模态大语言模型(LLM)摘要时出现关键信息遗漏与跨模态幻觉,这些问题源于长程依赖建模中的注意力漂移及跨模态对齐差距。为解决该问题,我们引入MMLDSum-Bench,这是一个覆盖多领域、上下文长度尺度及图文模态分布的高质量多模态长文档摘要基准。我们进一步提出MMLDSum-LLM,这是一个可复现的两阶段训练框架,结合带视觉对齐加权损失与关键词感知加权损失的监督微调,以及带多目标奖励(关键词覆盖率、图文对齐(ITA)、ROUGE、长度控制)的GRPO。在MMLDSum-Bench上的大量实验,在统一评估协议下对比主流闭源与开源多模态模型——包括LLM-as-a-judge评分、原子断言精确率/召回率、图文对齐(ITA)及ROUGE——表明我们的方法显著提升了关键信息覆盖率与跨模态一致性。

英文摘要

Multimodal long documents are core carriers of professional knowledge, where critical evidence is sparsely distributed across paragraphs and modalities. This easily causes key information omission and cross-modal hallucinations in summarization by multimodal LLMs. These issues stem from attention drift in long-range dependency modeling and gaps in inter-modal alignment. To address this, we introduce MMLDSum-Bench, a high-quality benchmark for multimodal long-document summarization, covering multiple domains, context-length scales, and visual-textual modality distributions. We further propose MMLDSum-LLM, a reproducible two-stage training framework that combines supervised fine-tuning with visual-alignment weighted loss and keyword-aware weighted loss, followed by GRPO with a multi-objective reward (keyword coverage, image-text alignment, ROUGE, and length control). Extensive experiments on MMLDSum-Bench, comparing against leading closed-source and open-source multimodal models under a unified evaluation protocol - including LLM-as-a-judge scoring, atomic-claim precision/recall, image-text alignment (ITA), and ROUGE - demonstrate that our approach significantly improves key-information coverage and cross-modal consistency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑