MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Comments 29 pages, 16 figures
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Comments 29 pages, 16 figures
专题命中 多模态评测 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments 28 pages, 14 figures, 5 tables. Evaluation code (LLM-as-a-judge and Markdown TEDS) is available at https://github.com/jinkhye/MyFinMarkdown. The development dataset and evaluation benchmark are available on Hugging Face at https://huggingface.co/datasets/jinkhye/MyFinMarkdown-sample and https://huggingface.co/datasets/jinkhye/MyFinMarkdown-bench respectively
机构 * School of Computer Science Fudan University(复旦大学计算机科学学院) ; Department of Computer Science University of Rochester(罗切斯特大学计算机科学系) ; School of Economics Fudan University(复旦大学经济学院)
专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments To appear in the International Conference on Computer Vision, ICCV 2025
专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI
专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 24 pages, 10 figures
专题命中 多模态评测 :multi-modal(abstract)
Journal ref IEEE Robotics and Automation Letters, vol. 10, no. 9, pp. 9312-9319, 2025