arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-31 至 2025-10-31 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. 文档图表理解 3 篇

2510.26339 2025-10-31 cs.CV cs.AI 76%

GLYPH-SR: Can We Achieve Both High-Quality Image Super-Resolution and High-Fidelity Text Recovery via VLM-guided Latent Diffusion Model?

Mingyu Sung, Seungjae Ham, Kangwoo Kim, Yeokyoung Yoon, Sangseok Yun, Il-Min Kim, Jae-Mo Kang

机构 * Department of Artificial Intelligence(人工智能系) Kyungpook National University(庆尚国立大学) Department of Electrical and Computer Engineering(电气电子工程系) Queen’s University(皇后大学) Department of Information and Communications Engineering(信息与通信工程系) Pukyong National University(浦项国立大学)

专题命中 文档图表理解 :VLM(title);分类 cs.CV、cs.AI

Comments 11 pages, 6 figures. Includes supplementary material. Under review as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21497 2025-10-31 cs.CV cs.AI cs.CL cs.MA 62%

Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers

Wei Pang, Kevin Qinghong Lin, Xiangru Jian, Xi He, Philip Torr

机构 * University of Waterloo(滑铁卢大学) University of Oxford(牛津大学) Vector Institute(向量研究所)

专题命中 文档图表理解 :VLM(abstract);分类 cs.CV、cs.AI

Comments Project Page: https://github.com/Paper2Poster/Paper2Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25160 2025-10-31 cs.CL cs.AI cs.IR 57%

Model-Document Protocol for AI Search

Hongjin Qian, Zheng Liu

机构 * BAAI(北京人工智能研究院)

专题命中 文档图表理解 :grounding(abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏