NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
机构 * IIT Kharagpur(印度理工学院Kharagpur分校) ; Ashoka University(阿什oka大学)
专题命中 视觉问答 :vision-language model(abstract);分类 cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * IIT Kharagpur(印度理工学院Kharagpur分校) ; Ashoka University(阿什oka大学)
专题命中 视觉问答 :vision-language model(abstract);分类 cs.AI
机构 * School of Artificial Intelligence and Software Engineering, Nanyang Normal University, Henan, China(人工智能与软件工程学院,南阳师范学院,河南) ; Institute for Artificial Intelligence, Peking University, Beijing, China(人工智能研究院,北京大学,北京) ; Collaborative Innovation Center of Intelligent Explosion-proof Equipment, Henan, China(智能防爆设备协同创新中心,河南)
专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV、cs.LG
Comments Accepted at the International Joint Conference on Artificial Intelligence (IJCAI 2025)
机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) ; Leibniz-Institut für Analytische Wissenschaften – ISAS – e.V.(莱比锡分析科学研究所(ISAS)) ; Department of Pathology, The Sixth Affiliated Hospital, Sun Yat-sen University(中山大学第六附属医院病理科部) ; Institute of Pathology, University Hospital Essen(埃森大学医院病理科研究所) ; Academy for Multidisciplinary Studies, Capital Normal University(首都师范大学多学科研究学院)
专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV
机构 * University of Bristol(布里斯托大学) ; University of Amsterdam(阿姆斯特丹大学)
专题命中 视觉推理 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
Comments 11 pages including references, 6 figures. Accepted at IWCS 2025
机构 * Vanderbilt University(范德比尔特大学) ; Weill Cornell Medicine(韦尔·科恩医学中心) ; Vanderbilt University Medical Center(范德比尔特大学医学中心) ; UT MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心)
专题命中 视觉推理 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV
机构 * Northeastern University(东北大学) ; Memorial Sloan Kettering Cancer Center(纪念斯隆凯特琳癌症中心)
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.LG
机构 * School of Informatics, Xiamen University(厦门大学信息学院) ; School of Computer Science, Nanjing University(南京大学计算机科学学院) ; School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) ; Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(教育部多媒体可信感知与高效计算重点实验室,厦门大学)
专题命中 视觉定位与Grounding :VLM(title,abstract);grounding(title,abstract);分类 cs.CV、cs.AI
机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
专题命中 视觉定位与Grounding :grounding(title,abstract);LLaVA(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
Comments Accepted by COLM 2025
机构 * University of Aberdeen(阿伯丁大学)
专题命中 视觉定位与Grounding :vision-language model(title);vision language model(abstract);VLM(abstract);分类 cs.CV
Comments Accepted to MICCAI 2025
机构 * University of Science and Technology of China(科学技术大学) ; iFLYTEK Research(iFLYTEK研究院)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Seoul National University(首尔国立大学) ; NVIDIA(NVIDIA公司)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments Preprint. Project page: https://jaeyeonkim99.github.io/wow_bench/
机构 * Department of Computer Science, University of Virginia(大学计算机科学系) ; University of Virginia School of Medicine(弗吉尼亚大学医学院) ; Virginia Polytechnic Institute and State University(弗吉尼亚理工学院和州立大学) ; Biocomplexity Institute and Initiative, University of Virginia(大学生物复杂性研究所) ; Division of Infectious Diseases & International Health, University of Virginia School of Medicine(大学感染性疾病与国际卫生分会)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG
机构 * Computer Vision Center, Universitat Autònoma de Barcelona, Spain(巴塞罗那自治大学计算机视觉中心)
专题命中 文档图表理解 :vision-language model(abstract);visual question answering(abstract);分类 cs.CV
Comments Accepted at Workshop on Machine Learning in Document Analysis and Recognition (ICDAR WML 2025), Wuhan, China
机构 * Alibaba Group(阿里巴巴集团) ; Beijing University of Posts and Telecommunications(北京邮电大学)
专题命中 文档图表理解 :VLM(abstract);visual language model(abstract);分类 cs.CV
Comments Accepted by CIKM 2025
专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
机构 * Johns Hopkins University(约翰霍普金斯大学) ; Meta AI Research(Meta AI 研究)
专题命中 VLM训练与架构 :visual language model(title,abstract);分类 cs.CV
Comments EMNLP 2025 (Findings)
机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学)
专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments BMVC 2025
机构 * ByteDance Inc.(字节跳动公司)
专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV
Comments project url: https://one-reward.github.io
机构 * The State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室) ; Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术区(滨江)区块链与数据安全研究院) ; Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) ; School of Computer Science, Beijing University of Posts and Telecommunications(北京邮电大学计算机学院)
专题命中 VLM训练与架构 :LLaVA(abstract);分类 cs.CV、cs.AI
Comments Accepted at EMNLP 2025 Main
专题命中 VLM训练与架构 :LLaVA(abstract);分类 cs.CV
Comments 10 pages, 3 figures
专题命中 VLM训练与架构 :vision language model(abstract);分类 cs.CV
Comments Paper accepted at the eXCV Workshop at ECCV 2024. Supplementary material included. Code available at https://github.com/idiap/itm
机构 * Interactive Technologies Institute and NOVA LINCS Faculty of Exact Sciences and Engineering University of Madeira Portugal(互动技术研究所和NOVA LINCS精确科学与工程学院马德拉大学) ; Department of Informatics University of Bergen Norway(信息学院卑尔根大学挪威) ; Valencian Research Institute for Artificial Intelligence Universitat Politècnica de València Spain(瓦伦西亚人工智能研究机构瓦伦西亚理工大学西班牙) ; Leverhulme Centre for the Future of Intelligence and Valencian Research Institute for Artificial Intelligence Spain(未来智能中心和瓦伦西亚人工智能研究机构西班牙)
专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV
Comments 54 pages (42 pages of appendix). Accepted for publication at the ECAI 2025 conference
机构 * Fudan University(复旦大学) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 其他VLM :multimodal large language model(abstract)