SimpleOCR: Rendering Visualized Questions to Teach MLLMs to Read
SimpleOCR: 将可视化问题呈现以教导大语言模型阅读
机构 * UNC-Chapel Hill(北卡罗来纳大学教堂山分校) ; Carnegie Mellon University(卡内基梅隆大学) ; University of Michigan(密歇根大学) ; Adobe Research(Adobe研究)
AI总结 SimpleOCR通过强制模型视觉参与,解决多模态大语言模型在图像中阅读文本的性能问题,提升模型的视觉文本提取能力。