arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-21 至 2025-10-21 共收录 66 信号源:cs.CV, cs.AI, cs.LG

1. VLM训练与架构 13 篇

2503.01879 2025-10-21 cs.MM cs.CV cs.SD eess.AS 57%

Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision

Che Liu, Yingji Zhang, Dong Zhang, Weijie Zhang, Chenggong Gong, Yu Lu, Shilin Zhou, Ziliang Gan, Ziao Wang, Haipang Wu, Ji Liu, André Freitas, Qifan Wang, Zenglin Xu, Rongjuncheng Zhang, Yong Dai

机构 * Imperial College London(伦敦帝国学院) University of Manchester(曼彻斯特大学) HiThink Research(HiThink研究院) Soochow University(苏州大学) Hong Kong Baptist University(香港 Baptist大学) Idiap Research Institute(Idiap研究 institute) Meta AI Fudan University(复旦大学)

专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV

Comments Project: https://github.com/HiThink-Research/NEXUS-O

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15682 2025-10-21 cs.CV 57%

ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval

Guanqi Zhan, Yuanpei Liu, Kai Han, Weidi Xie, Andrew Zisserman

机构 * VGG, University of Oxford(牛津大学视觉几何组) The University of Hong Kong(香港大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV

Comments Accepted by CBMI 2025 (IEEE International Conference on Content-Based Multimedia Indexing)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12943 2025-10-21 cs.CL 50%

The Curious Case of Curiosity across Human Cultures and LLMs

Angana Borah, Zhijing Jin, Rada Mihalcea

机构 * University of Michigan - Ann Arbor(密歇根大学安阿伯分校) University of Toronto(多伦多大学) Vector Institute(向量研究所) MPI for Intelligent Systems, Tubingen, Germany(图宾根德国智能系统研究所)

专题命中 VLM训练与架构 :grounding(abstract)

Comments Preprint (Paper under review)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他VLM 3 篇

2510.17002 2025-10-21 cs.LG 70%

EEschematic: Multimodal-LLM Based AI Agent for Schematic Generation of Analog Circuit

Chang Liu, Danial Chitnis

机构 * School of Engineering The University of Edinburgh Edinburgh, UK(工程学院 苏格兰爱丁堡大学)

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16448 2025-10-21 cs.LG cs.AI 62%

Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts

Yongxiang Hua, Haoyu Cao, Zhou Tao, Bocheng Li, Zihao Wu, Chaohu Liu, Linli Xu

机构 * University of Science and Technology of China(科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI、cs.LG

Comments ACM MM25

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16034 2025-10-21 cs.CV 57%

VisualLens: Personalization through Task-Agnostic Visual History

Wang Bill Zhu, Deqing Fu, Kai Sun, Yi Lu, Zhaojiang Lin, Seungwhan Moon, Kanika Narang, Mustafa Canim, Yue Liu, Anuj Kumar, Xin Luna Dong

机构 * Meta University of Southern California(南加州大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏