Harnessing PDF Data for Improving Japanese Large Multimodal Models
利用PDF数据提升日语大规模多模态模型
机构 * The University of Tokyo(东京大学) ; National Institute of Informatics(信息处理研究所)
专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 本文通过利用日本PDF数据提升日语大规模多模态模型的性能,采用自动化流程提取图像-文本对并构建指令数据,实验结果显示在Heron-Bench上性能提升达2.1%-13.8%。
Comments Accepted to ACL2025 Findings. Code: https://github.com/ku21fan/PDF-JLMM
Journal ref Findings of the Association for Computational Linguistics: ACL 2025