arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-02-13 至 2026-02-13 共收录 2 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 2 篇

2504.04988 2026-02-13 cs.CV cs.AI 73%

Remote Sensing Retrieval-Augmented Generation: Bridging Remote Sensing Imagery and Comprehensive Knowledge with a Multi-Modal Dataset and Retrieval-Augmented Generation Model

遥感检索增强生成:通过多模态数据集和检索增强生成模型连接遥感图像与综合知识

Congcong Wen, Yiting Lin, Xiaokang Qu, Nan Li, Yong Liao, Xiang Li, Hui Lin

机构 * School of Cyber Science and Technology, University of Science and Technology of China(信息科学技术学院,中国科学技术大学) China Academy of Electronics and Information Technology(电子信息技术研究院)

专题命中 视觉问答 :VLM(abstract);visual question answering(abstract);分类 cs.CV、cs.AI

AI总结 本文提出RS-RAG框架,通过多模态数据集和检索增强生成模型,提升遥感图像与综合知识的连接能力,有效提升复杂查询的语义推理性能。

Comments Accepted by IEEE Geoscience and Remote Sensing Magazine (GRSM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12203 2026-02-13 cs.CL 67%

ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images

ExStrucTiny:一种用于从文档图像中进行模式变量结构信息提取的基准

Mathieu Sibue, Andres Muñoz Garza, Samuel Mensah, Pranav Shetty, Zhiqiang Ma, Xiaomo Liu, Manuela Veloso

机构 * J.P. Morgan AI Research(摩根大通人工智能研究)

专题命中 视觉问答 :vision language model(abstract);visual question answering(abstract)

AI总结 ExStrucTiny是一个新的文档图像结构信息提取基准,通过结合手动和合成样本,涵盖更多多样化的文档类型和提取场景,旨在提升通用模型在结构化信息提取中的性能。

Comments EACL 2026, main conference

详情

展开后加载摘要…

URL PDF HTML 收藏