MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware
MMLDSum-LLM:结合视觉对齐与关键词感知的多模态长文档摘要方法
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.AI
AI总结 针对多模态长文档摘要的关键信息遗漏与跨模态幻觉问题,提出结合视觉对齐与关键词感知的两阶段训练框架MMLDSum-LLM,在自研基准MMLDSum-Bench上验证了其性能优势。