Distilling Multilingual Vision-Language Models: When Smaller Models Stay Multilingual
机构 * VISTEC ; AI Singapore
专题命中 视觉问答 :vision-language model(title,abstract)
Comments Work in progress
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * VISTEC ; AI Singapore
专题命中 视觉问答 :vision-language model(title,abstract)
Comments Work in progress
机构 * Columbia University, United States Shanghai Jiao Tong University, China San Francisco State University, United States Carnegie Mellon University, United States Texas A\&M University, College Station, United States University of California, San Diego, United States
专题命中 视觉问答 :LLaVA(abstract);visual question answering(abstract);分类 cs.CV
机构 * Leibniz University Hannover(莱比锡大学汉诺威分校) ; L3S Research Center(L3S研究中心) ; Microsoft(微软公司)
专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV
Comments Accepted to NeurIPS 2025
机构 * NVIDIA ; NIH/NCI ; CHOP/UPenn ; Basaksehir Cam and Sakura City Hospital
专题命中 视觉推理 :visual language model(title);vision-language model(abstract);分类 cs.CV
Comments NV-Reason-CXR-3B
机构 * Rochester Institute of Technology(罗切斯特技术研究所) ; Snap Inc.(Snap公司) ; University of Rochester(罗切斯特大学) ; DEVCOM Army Research Laboratory(陆军研究实验室)
专题命中 视觉推理 :visual reasoning(title);vision-language model(abstract);分类 cs.AI
Comments NeurIPS 2025
机构 * Tsinghua University(清华大学) ; Pengcheng Laboratory(鹏城实验室) ; Sun Yat-sen University(中山大学)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV
机构 * Fudan University(复旦大学) ; Meituan(美团) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 视觉推理 :vision-language model(abstract);visual reasoning(abstract);分类 cs.CV、cs.AI、cs.LG
Comments Preprint
机构 * CUHK(香港中文大学) ; IMIXR ; MMLab ; Peking University(北京大学) ; Northeastern University(东北大学)
专题命中 视觉推理 :visual reasoning(abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments Project Page: https://video-cof.github.io
机构 * Indian Institute of Information Technology Dharwad, India(印度达拉瓦德信息科技学院) ; Indian Institute of Technology Indore, India(印度印度理工学院) ; Malaviya National Institute of Technology Jaipur, India(马拉维亚国家理工学院)
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments Accepted at IEEE International Conference on Data Mining (ICDM) 2025
机构 * Humains AI Research(Humains人工智能研究) ; Inpris Ltd(Inpris公司)
专题命中 视觉推理 :grounding(abstract);分类 cs.AI、cs.LG
机构 * Electronics and Telecommunications Research Institute, Republic of Korea(韩国电子电信研究院) ; POSTECH ; Sungkyunkwan University(全南大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI、cs.LG
Comments NeurIPS 2025, 38 pages, 8 figures
机构 * University of Technology Sydney(技术大学悉尼大学) ; Tongji University(同济大学) ; University of Melbourne(墨尔本大学) ; University of Sydney(悉尼大学) ; Macquarie University(麦考瑞大学)
专题命中 视觉推理 :MLLM(abstract);分类 cs.AI、cs.LG
Comments Accepted at NeurIPS 2025
机构 * Case Western Reserve University(凯斯西储大学)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.LG
Comments Datasets link: https://huggingface.co/datasets/LLDDSS/Causal3D_Dataset
机构 * Xiamen University(厦门大学) ; Anyang Normal University(安阳师范学院) ; Tencent YouTu Lab(腾讯YouTu实验室) ; Tencent SSV(腾讯SSV)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
机构 * School of Informatics, University of Edinburgh(信息学院,爱丁堡大学)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.LG
Comments Accepted to NeurIPS 2025
机构 * Department of Medicine I, LMU University Hospital, LMU Munich, Germany(慕尼黑大学医学部第一部门,LMU大学医院,慕尼黑,德国) ; Lunit Inc.(Lunit公司)
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
Comments Accepted to Computer Vision for Automated Medical Diagnosis (CVAMD) Workshop at ICCV 2025
机构 * EPFL(瑞士联邦理工学院) ; MILA(蒙特利尔人工智能研究院)
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV
Journal ref 2025 Conference on Empirical Methods in Natural Language Processing
机构 * Stanford University(斯坦福大学) ; University of Toronto(多伦多大学) ; University of Pennsylvania(宾夕法尼亚大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG
专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV
机构 * Key Laboratory of Multimedia Trusted Perception(多媒体可信感知关键实验室) ; Efficient Computing, Ministry of Education of China, Xiamen University(高效计算、教育部中国 ministry of education、厦门大学) ; Institute of Artificial Intelligence, Xiamen University(人工智能研究院、厦门大学) ; Peng Cheng Laboratory, Shenzhen, China(鹏城实验室、深圳中国)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG
机构 * Tencent Youtu Lab(腾讯优图实验室)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments 12 pages, 7 figures
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV
Comments EMNLP 2025 Oral; Project Homepage: https://yanzehong.github.io/trust-vl/
专题命中 视觉定位与Grounding :grounding(abstract)
机构 * Department of Artificial Intelligence(人工智能系) ; Kyungpook National University(庆尚国立大学) ; Department of Electrical and Computer Engineering(电气电子工程系) ; Queen’s University(皇后大学) ; Department of Information and Communications Engineering(信息与通信工程系) ; Pukyong National University(浦项国立大学)
专题命中 文档图表理解 :VLM(title);分类 cs.CV、cs.AI
Comments 11 pages, 6 figures. Includes supplementary material. Under review as a conference paper at ICLR 2026
机构 * University of Waterloo(滑铁卢大学) ; University of Oxford(牛津大学) ; Vector Institute(向量研究所)
专题命中 文档图表理解 :VLM(abstract);分类 cs.CV、cs.AI
Comments Project Page: https://github.com/Paper2Poster/Paper2Poster
机构 * BAAI(北京人工智能研究院)
专题命中 文档图表理解 :grounding(abstract);分类 cs.AI
Comments 10 pages
专题命中 GUI与屏幕智能体 :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments 39 pages, 24 figures
机构 * University of California Santa Cruz(加州大学圣克ruz分校) ; Northeastern University(东北大学) ; Accenture(Accenture公司)
专题命中 GUI与屏幕智能体 :vision-language model(abstract);分类 cs.CV、cs.AI
Comments NeurIPS 2025
机构 * Department of Computer Science(计算机科学系) ; Virginia Tech(弗吉尼亚理工学院)
专题命中 VLM训练与架构 :vision language model(title);vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.LG
机构 * KAIST(韩国科学技术院)
专题命中 VLM训练与架构 :VLM(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI
Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: The First Workshop on Generative and Protective AI for Content Creation