arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-08-11 至 2025-08-11 共收录 2 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 2 篇

2505.19015 2025-08-11 cs.CV cs.MM 83%

Can Multimodal Large Language Models Understand Spatial Relations?

Jingping Liu, Ziyan Liu, Zhedong Cen, Yan Zhou, Yinan Zou, Weiyan Zhang, Haiyun Jiang, Tong Ruan

机构 * School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China(信息科学与工程学院,东华大学,上海,中国) School of Computer Science, Fudan University, Shanghai, China(计算机科学学院,复旦大学,上海,中国)

专题命中 视觉推理 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

Comments 13 pages, 7 figures, published to ACL 2025

Journal ref In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 620-632, 2025, Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06009 2025-08-11 cs.CV 79%

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models

Jun Feng, Zixin Wang, Zhentao Zhang, Yue Guo, Zhihan Zhou, Xiuyi Chen, Zhenyang Li, Dawei Yin

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV

Comments 29 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏