ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.LG
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.LG
机构 * School of Cyber Science and Engineering, Qufu Normal University(网络科学与工程学院,曲阜师范大学) ; Department of Computer Science, University of Verona(计算机科学系,威尼斯大学)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
机构 * School of Information Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)信息科学与技术学院) ; Guangdong Provincial Key Laboratory of Space-Aerial Networking and Intelligent Sensing(广东省空间-航空网络与智能感知重点实验室) ; Information Systems Technology and Design, Singapore University of Technology and Design(新加坡科技设计大学信息系统技术与设计) ; School of System Design and Intelligent Manufacturing, Southern University of Science and Technology(南方科技大学系统设计与智能制造学院)
专题命中 视觉推理 :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.LG
Comments The paper has been submitted to IEEE Internet of Things Magazine
机构 * Fujian University of Technology(福建工程学院) ; Fujian Provincial Key Laboratory of Big Data Mining and Applications(福建省大数据挖掘与应用重点实验室) ; Key Laboratory of Biomedical Imaging Science and System, Chinese Academy of Sciences(生物医学成像科学与系统重点实验室,中国科学院)
专题命中 视觉推理 :vision-language model(abstract);LLaVA(abstract);grounding(abstract);分类 cs.AI
机构 * Carnegie Mellon University(卡内基梅隆大学) ; VIT Bhopal(维捷商学院)
专题命中 视觉推理 :vision-language model(abstract);vision language model(abstract);visual reasoning(abstract);分类 cs.CV
Comments 4 pages, 6 figures, International Conference on Computer Vision, ICCV 2025
机构 * Keye Team, Kuaishou Group(快手集团Keye团队)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV
Comments Github page: https://github.com/Kwai-Keye/Keye
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Lab(上海人工智能实验室) ; Wuhan University(武汉大学) ; East China Normal Unversity(华东师范大学) ; Macao Polytechnic University(澳门 polytechnic university) ; Southern University of Science and Technology(南方科技大学)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
Comments 6 pages, 6 figures
机构 * Beihua University(白华大学) ; Kasem Bundit University(Kasem Bundit大学)
专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV
机构 * Meta FAIR
专题命中 视觉推理 :VLM(abstract);分类 cs.AI
机构 * Northwestern University(西北大学) ; University of Illinois at Chicago(伊利诺伊大学香槟分校)
专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI
Comments Accepted to the Trustworthy FMs workshop in ICCV 2025
机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) ; Ministry of Education Key Laboratory of Intelligent Networks and Network Security(教育部智能网络与网络安全重点实验室) ; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering(陕西省大数据知识工程重点实验室) ; IHPC, Agency for Science, Technology and Research(科技研究局IHPC) ; College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV
机构 * College of Electronics and Information, Hangzhou Dianzi University(电子信息学院,杭州电子大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
机构 * University of Virginia(弗吉尼亚大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
Comments Conference on Robot Learning (CoRL) 2025. Project website: https://o3afford.github.io/
机构 * Apple(苹果公司)
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.LG
Comments ICCV 2025
机构 * College of Computing and Data Science, Nanyang Technological University(computing and Data Science学院,南洋理工大学) ; School of Electrical and Electronic Engineering, Nanyang Technological University(Electrical and Electronic Engineering学院,南洋理工大学) ; School of Artificial Intelligence, Wuhan University(Artificial Intelligence学院,武汉大学) ; ByteDance(字节跳动)
专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Extended version of our conference paper arXiv:2410.09855
机构 * Sharif University of Technology(谢里夫理工大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG
Comments Accepted to TMLR 2025
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI
机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
专题命中 视觉定位与Grounding :MLLM(abstract);分类 cs.CV
机构 * Centre for AI and ML, Edith Cowan University(人工智能与机器学习中心,埃德温·科温大学) ; University of Manchester(曼彻斯特大学)
专题命中 视觉定位与Grounding :vision-language model(abstract)
Comments 15 pages, 4 figures. Preprint
专题命中 视觉定位与Grounding :grounding(abstract)
机构 * China Academy of Information and Communications Technology(中国信息通信技术研究院) ; Renmin University of China(中国人民大学) ; University of International Business and Economics(国际经济贸易大学)
专题命中 视觉定位与Grounding :grounding(abstract)
机构 * NetEase, Inc.(网易公司) ; Beihang University(北京航空航天大学) ; Tsinghua University(清华大学) ; City University of Hong Kong(香港城市大学) ; Independent Researcher & Technical Artists(独立研究者及技术艺术家)
专题命中 文档图表理解 :multimodal large language model(title);分类 cs.CV、cs.AI、cs.LG
机构 * Department of Mechanical and Aerospace Engineering, Tandon School of Engineering, New York University(机械与航空航天工程系,坦顿工程学院,纽约大学) ; GenAuto.ai by General Autonomy Inc.(General Autonomy Inc. 的 GenAuto.ai)
专题命中 GUI与屏幕智能体 :vision-language model(title,abstract)
机构 * Tianjin Agricultural University(天津农业大学) ; Walailak University(Walailak大学)
专题命中 GUI与屏幕智能体 :LLaVA(abstract);分类 cs.CV
专题命中 幻觉与鲁棒性 :vision language model(title,abstract);VLM(title,abstract);分类 cs.CV、cs.AI
机构 * University of Southern California(南加州大学) ; University of Bristol(布里斯托大学) ; Centre National de la Recherche Scientifique(法国国家科学研究中心) ; University of California, Los Angeles(加州大学洛杉矶分校)
专题命中 幻觉与鲁棒性 :vision language model(title);vision-language model(abstract);LLaVA(abstract);分类 cs.CV
Comments Accepted to EMNLP 2025 (Main)
机构 * School of Information and Communication Engineering, Dalian University of Technology(信息与通信工程学院,大连理工大学) ; School of Computer Science and Technology, Anhui University(计算机科学与技术学院,安徽大学) ; New Laboratory of Pattern Recognition (NLPR) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS) Institute of Automation, Chinese Academy of Sciences (CASIA)(模式识别新实验室(NLPR)多模态人工智能系统国家重点实验室(MAIS)自动化研究所,中国科学院(CASIA))
专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.AI、cs.LG
机构 * Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, Nanjing, China(新型一代人工智能技术及其交叉应用重点实验室(东南大学),教育部,南京,中国) ; School of Biomedical Engineering, Tsinghua University(生物医学工程学院,清华大学) ; Department of Biomedical Engineering and Department of Electrical and Computer Engineering, National University of Singapore(生物医学工程系和电子与计算机工程系,新加坡国立大学) ; School of Information Science and Engineering, Southeast University(信息科学与工程学院,东南大学) ; Department of Biostatistics, Center for Global Health, School of Public Health, Nanjing Medical University(流行病学系,全球健康中心,公共卫生学院,南京医科大学) ; Institute of High-Performance Computing, Agency for Science, Technology and Research, Singapore(高性能计算研究所,科技研究局,新加坡)
专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV
Comments 20 pages, 8 figures
机构 * Northwestern University(西北大学) ; University of Illinois at Chicago(伊利诺伊大学香槟分校)
专题命中 幻觉与鲁棒性 :multimodal large language model(abstract);分类 cs.AI
Comments accepted by the Trustworthy FMs workshop in ICCV 2025
专题命中 幻觉与鲁棒性 :vision-language model(abstract)
Comments 7 pages, 3 figures, submitted to EMNLP 2025 and ECAT Research Workshop 2025