arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-10 至 2025-09-10 共收录 2 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 2 篇

2509.07846 2025-09-10 cs.AI 57%

Aligning LLMs for the Classroom with Knowledge-Based Retrieval -- A Comparative RAG Study

Amay Jain, Liu Cui, Si Chen

机构 * Student(学生) Downingtown STEM Academy Department of Computer Science(计算机科学系) West Chester University of Pennsylvania(宾夕法尼亚州韦斯特切斯特大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17471 2025-09-10 cs.CL 50%

FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain

Suifeng Zhao, Zhuoran Jin, Sujian Li, Jun Gao

机构 * Key Laboratory of High Confidence Software Technologies, CS, Peking University, China(北京大学高可信软件技术重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Multimedia Information Processing, School of Computer Sciences, Peking University(北京大学多媒体信息处理国家重点实验室)

专题命中 视觉问答 :multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏