LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
LogicOCR: 大型多模态模型在文本丰富的图像上逻辑推理是否表现优异?
机构 * School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University(计算机学院、多媒体软件国家工程研究中心、人工智能研究院、多媒体与网络通信工程湖北省重点实验室、武汉大学)
专题命中 逻辑推理 :reasoning(title,abstract);logical reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract)
AI总结 LogicOCR通过构建包含生成和现实图像问题的基准测试,评估大型多模态模型在文本丰富图像上的逻辑推理能力,并提出TextCue方法提升模型对关键文本区域的感知。
Comments GitHub: https://github.com/MiliLab/LogicOCR