Measuring How (Not Just Whether) VLMs Build Common Ground
机构 * Northeastern University(东北大学)
专题命中 视觉问答 :vision language model(abstract);VLM(abstract);grounding(abstract);分类 cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Northeastern University(东北大学)
专题命中 视觉问答 :vision language model(abstract);VLM(abstract);grounding(abstract);分类 cs.AI
专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract)
机构 * Bradley Department of Electrical and Computer Engineering(电气与计算机工程系)
专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.LG
Comments Accepted at IEEE GLOBECOM 2025
机构 * School of Computer Science(计算机科学学院) ; University College Dublin(都柏林大学) ; School of Computing(计算机科学学院) ; Dublin City University(都柏林城市大学) ; School of Electronic Engineering(电子工程学院) ; Trinity College Dublin(都柏林三一学院)
专题命中 视觉推理 :VLM(abstract);分类 cs.CV、cs.AI、cs.LG
机构 * Wanfu Wang, Qipeng Huang, Guangquan Xue, Xiaobo Liang, Juntao Li(作者)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision language model(abstract);分类 cs.CV、cs.AI
机构 * Department of Biomedical Engineering(生物医学工程系) ; Georgia Institute of Technology(佐治亚理工学院) ; Department of Machine Learning(机器学习系) ; Department of Radiation Oncology(放射肿瘤科) ; Emory University School of Medicine(埃默里大学医学院) ; University of Southern California(南加州大学)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);visual question answering(abstract);分类 cs.CV
机构 * Xi’an Jiaotong University(西安交通大学) ; SGIT AI Lab(SGIT人工智能实验室) ; Zhejiang University of Technology(浙江工业大学) ; Huawei(华为)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI
机构 * College of Electronics and Information Engineering, Wuyi University(威怡大学电子与信息工程学院) ; School of Electronic and Information Engineering and the Key Laboratory of Big Data and Intelligent Robot, Ministry of Education, South China University of Technology(电子与信息工程学院和大数据与智能机器人重点实验室,华南理工大学) ; State Key Laboratory of Lunar and Planetary Sciences, Macau University of Science and Technology(澳门大学地球和行星科学国家重点实验室) ; College of Engineering and Computer Science, California State University, Northridge(工程与计算机科学学院,加州大学北岭分校) ; Department of Geography, The University of Hong Kong(地理系,香港大学) ; Faculty of Computer Science and Engineering, S(计算机科学与工程学院,S)
专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
机构 * Politecnico di Torino(托斯尼亚理工学院)
专题命中 视觉定位与Grounding :visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV
Comments Accepted at ICIAP 2025
机构 * TOELT LLC AI lab(TOELT LLC人工智能实验室)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
机构 * PES University(PES大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(abstract)
机构 * Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系) ; School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院) ; Idiap Research Institute(Idiap研究机构) ; National Biomarker Centre, CRUK Manchester Institute(国家生物标志物中心、CRUK曼彻斯特研究所)
专题命中 视觉定位与Grounding :grounding(abstract)
Comments EMNLP 2025 Camera-Ready Version
机构 * Sprinklr
专题命中 文档图表理解 :vision-language model(abstract);InternVL(abstract);分类 cs.AI
Comments Sprinklr OCR provides a fast and compute light way of performing OCR
机构 * Zhejiang Key Lab of Accessible Perception \& Intelligent Systems, Zhejiang University Hangzhou China ; Zhejiang University
专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);分类 cs.AI
Comments Paper accepted to ACM MM 2025
机构 * Zalando SE Berlin Germany(泽尔安多德国分公司) ; Zalando Switzerland AG Zürich Switzerland(泽尔安多瑞士分公司)
专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.LG
机构 * Department of Data Science & AI(数据科学与人工智能系) ; Department of Electrical and Computer Systems Engineering(电气与计算机系统工程系)
专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV
专题命中 幻觉与鲁棒性 :vision language model(abstract)
机构 * University of Waterloo(滑铁卢大学) ; Toronto Metropolitan University(多伦多 Metropolitan 大学)
专题命中 幻觉与鲁棒性 :grounding(abstract)
机构 * Sungkyunkwan University(顺天大学) ; Deakin University(德金大学)
专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV
Comments ICCV 2025 - LIMIT Workshop
机构 * Flowers AI & CogSci Lab, Inria, France(Flowers AI与认知科学实验室,Inria,法国) ; MIT, USA(麻省理工学院,美国)
专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.AI
机构 * University of Southhampton(南安普顿大学) ; KAIST(韩国科学技术院)
专题命中 VLM训练与架构 :VLM(abstract);分类 cs.CV、cs.AI
Comments 17 pages, 5 figures, 9 tables
机构 * Stanford University(斯坦福大学)
专题命中 其他VLM :vision-language model(abstract);vision language model(abstract);分类 cs.CV
Comments Accepted at ICCV 2025. Project website: https://dense-functional-correspondence.github.io/
机构 * Lu Wang(卢王) ; Hao Chen(何晨) ; Siyu Wu(武士) ; Zhiyue Wu(吴致岳) ; Hao Zhou(周浩) ; Chengfeng Zhang(张成峰) ; Ting Wang(王婷) ; Haodi Zhang(张浩迪)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI、cs.LG
机构 * College of Control Science and Engineering, Zhejiang University, China(控制科学与工程学院,浙江大学,中国) ; Huzhou Institute of Industrial Control Technology, China(湖州工业控制技术研究所,中国) ; School of Computer Science and Engineering, Nanyang Technological University, Singapore(计算机科学与工程学院,南洋理工大学,新加坡) ; School of Computer Science, Shanghai Jiao Tong University, China(计算机科学学院,上海交通大学,中国) ; School of Computing, National University of Singapore, Singapore(计算学院,新加坡国立大学,新加坡) ; Research, Singapore(研究,新加坡)
专题命中 其他VLM :vision language model(abstract);分类 cs.CV、cs.AI
Comments Accepted to ICML 2025
机构 * College of Engineering & Advanced Computing(工程与高级计算学院)
专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV