Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
机构 * TCS-Research(TCS研究)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
Comments 10 pages, 11 figures, 6 tables
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * TCS-Research(TCS研究)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
Comments 10 pages, 11 figures, 6 tables
机构 * Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine, Zhejiang University, Hangzhou(牙科医院,口腔医学院,浙江大学医学院,浙江大学,杭州) ; College of Computer Science and Technology, Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University, Hangzhou(计算机科学与技术学院,浙江大学-伊利诺伊大学 Urbana-Champaign 院,浙江大学,杭州) ; Department of Orthodontics, Shanghai Ninth People’s Hospital, College of Stomatology, Shanghai Jiao Tong University School of Medicine(正畸科,上海第九人民医院,口腔医学院,上海交通大学医学院) ; Angelalign Technology Inc.(Angelalign 技术公司)
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV、cs.AI
机构 * School of Computer Science, Hubei University(湖北大学计算机学院)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
Comments 14 pages, 6 figures. ACCEPTED for publication as a REGULAR paper in the IEEE Transactions on Multimedia 2025
机构 * School of Computer, Electronics and Information(计算机、电子与信息学院) ; Guangxi University(广西大学) ; Guangxi Key Laboratory of Multimedia Communications and Network Technology(广西多媒体通信与网络技术重点实验室)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
Comments Accepted by the IEEE International Conference on Multimedia and Expo (ICME 2025) for oral presentation. © 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.LG
机构 * Banking Academy of Vietnam(越南银行学院) ; Dublin City University(都柏林城市大学)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
机构 * Stanford University(斯坦福大学)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.LG
机构 * National Yang Ming Chiao Tung University(阳明大学)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
机构 * Amity Centre for Artificial Intelligence, Amity University, Noida(阿米蒂人工智能中心,阿米蒂大学,诺伊达)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
Comments 8 pages, 4 figures Submitted to AAAI 26
机构 * SimulaMet - Simula Metropolitan Center for Digital Engineering, Oslo, Norway(SimulaMet - Simula Metropolitan Center for Digital Engineering,挪威奥斯陆) ; Simula Research Laboratory, Oslo, Norway(Simula研究实验室,挪威奥斯陆) ; OsloMet - Oslo Metropolitan University, Oslo, Norway(OsloMet - 奥斯陆 Metropolitan 大学,挪威奥斯陆)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
机构 * University of Utah(犹他大学) ; University of Michigan(密歇根大学)
专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.AI、cs.LG
机构 * Information Sciences Institute(信息科学研究所) ; Institute for Creative Technologies(创意技术研究所) ; Department of Computer Science(计算机科学系) ; University of Southern California(南加州大学) ; University of California, Santa Barbara(加州大学圣巴巴拉分校) ; Aristotle University of Thessaloniki(希腊雅典纳大学)
专题命中 视觉问答 :vision language model(title);vision-language model(abstract);分类 cs.CV、cs.AI
Comments ACL 2025 Findings
机构 * LTCI, Télécom-Paris, Institut Polytechnique de Paris(LTCI, Télécom-Paris, Institut Polytechnique de Paris) ; Inria, LJK, Univ. Grenoble Alpes(Inria, LJK, 火车头 Grenoble Alpes) ; Universitat Autónoma de Barcelona(Autonomous University of Barcelona)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
Comments ICCV 2025, 8 pages. Code: https://github.com/IemProg/QUAD
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Abridge ; Mistral AI ; Johns Hopkins University(约翰霍普金斯大学)
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.AI、cs.LG
Comments Extended version of EMNLP 2024 paper arXiv:2411.04118. Includes additional results on clinical note QA tasks and supervised fine-tuning evaluations
机构 * University of Washington(华盛顿大学) ; Stanford University(斯坦福大学) ; University of Michigan(密歇根大学)
专题命中 视觉问答 :vision-language model(title);vision language model(abstract);分类 cs.CV、cs.LG
机构 * Beijing Institute of Technology(北京理工大学) ; Harbin Institute of Technology(哈尔滨工业大学) ; Guangdong University of Technology(广东工业大学)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
Comments Accepted at IJCAI 2025
机构 * Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)(迈德布兹泽德人工智能大学)
专题命中 视觉问答 :vision-language model(title);visual question answering(abstract);分类 cs.CV、cs.AI
Journal ref JIOT.2025.3579032
机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology, China(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学,中国) ; Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University, China(广东机器感知与智能计算实验室,深圳MSU-BIT大学,中国)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
Comments Accepted by IJCAI 2025
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
机构 * Tandon School of Engineering, New York University(纽约大学工程学院) ; School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电子与电气工程学院) ; College of Electronic and Information Engineering, Tongji University(同济大学电子与信息工程学院)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
Comments To be published in 2025 International Joint Conference on Neural Networks (IJCNN)
机构 * Key Laboratory of Complex Systems Modeling and Simulation, the School of Computer Science, Hangzhou Dianzi University(复杂系统建模与仿真重点实验室、计算机科学学院、杭州电子大学) ; HDU-ITMO Joint Institute, Hangzhou Dianzi University(杭州电子大学-ITMO联合学院) ; School of Computer Science and Information Engineering, Hefei University of Technology(计算机科学与信息工程学院、合肥工业大学) ; School of Intelligence Science and Engineering, Harbin Institute of Technology, Shenzhen(智能科学与工程学院、哈尔滨工业大学深圳校区)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.LG
Comments An extended journal version of our CVPR 2023 paper, which has been accepted at IEEE T-PAMI 2025. The original conference version can be referred to as the v1 version
专题命中 视觉问答 :visual question answering(title);LLaVA(abstract);分类 cs.CV、cs.AI
Journal ref Proceedings of Machine Learning Research, Volume 259, 2025, pp. 735--746
机构 * School of Artificial Intelligence, Tiangong University(天津工业大学人工智能学院) ; Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) ; Linying Medical Technology (Shenzhen) Co., Ltd.(临影医疗科技(深圳)有限公司)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI
机构 * Swinburne University of Technology(斯威本科技大学) ; PropelHealthAI
专题命中 视觉问答 :vision-language model(title);visual question answering(abstract);分类 cs.CV、cs.LG
Comments 10 pages, 4 figures
机构 * College of Electrical and Information Engineering, Hunan University(湖南大学电气与信息工程学院)
专题命中 视觉问答 :grounding(title,abstract);分类 cs.CV、cs.AI
Comments 8 pages, 6 figures, 3 tables
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Tongji University(同济大学)
专题命中 视觉问答 :visual language model(title);vision-language model(abstract);分类 cs.CV、cs.AI
机构 * University of Notre Dame(圣母大学) ; MBZUAI(穆罕默德·本·扎耶德人工智能大学) ; University of Southern California(南加州大学) ; University of Maryland, College Park(马里兰大学学院公园分校) ; KAUST(阿卜杜拉国王科技大学)
专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV、cs.AI
机构 * The Pennsylvania State University(宾夕法尼亚州立大学) ; GE Healthcare(GE医疗)
专题命中 视觉问答 :vision-language model(title);vision language model(abstract);分类 cs.CV、cs.AI
Comments Preprint, under review
机构 * Pattern Recognition Center, WeChat AI, Tencent Inc(腾讯公司微信AI模式识别中心)
专题命中 视觉问答 :vision language model(title);vision-language model(abstract);分类 cs.CV、cs.AI
机构 * KTH Royal Institute of Technology(瑞典皇家理工学院) ; University of Amsterdam(阿姆斯特丹大学) ; Digital Futures(数字未来机构) ; Helmholtz AI(亥姆霍兹人工智能研究所) ; Technical University of Munich(慕尼黑工业大学) ; MPI for Intelligent Systems, Tübingen(蒂宾根马克斯·普朗克智能系统研究所)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.LG
Comments Published at ICLR 2025