Is ChatGPT-5 Ready for Mammogram VQA?
机构 * Department of Radiation Oncology, Winship Cancer Institute, Emory University School of Medicine(放射肿瘤科、Winship癌症研究所、埃默里大学医学院)
专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
机构 * Department of Radiation Oncology, Winship Cancer Institute, Emory University School of Medicine(放射肿瘤科、Winship癌症研究所、埃默里大学医学院)
专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI
专题命中 视觉问答 :grounding(abstract);分类 cs.CV
Comments 21 pages, 6 figures, 17 tables
机构 * Research Institute of Electronic Science and Technology, University of Electronic Science and Technology of China(电子科学与技术研究院,电子科技大学) ; School of Aeronautics and Astronautics, University of Electronic Science and Technology of China(航空航天学院,电子科技大学)
专题命中 视觉推理 :vision-language model(title,abstract);visual reasoning(title,abstract);VLM(abstract);visual question answering(abstract)
专题命中 视觉推理 :vision language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI
机构 * Carnegie Mellon University, Robotics Institute(卡内基梅隆大学,机器人研究所)
专题命中 视觉推理 :grounding(title,abstract);分类 cs.CV、cs.AI
Comments 8 pages, 6 figures, published in IROS 2025
机构 * School of Automation Science and Engineering, Xi’an Jiaotong University(自动化科学与工程学院,西安交通大学) ; Interdisciplinary Graduate Programme, Nanyang Technological University(跨学科研究生项目,南洋理工大学) ; Lenovo Research, Lenovo(联想研究院,联想) ; School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) ; College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)
专题命中 视觉推理 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments 21 pages, 361 references
机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) ; School of Computer Science and Technology, Shandong Jianzhu University(山东建筑大学计算机科学与技术学院) ; School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) ; Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);分类 cs.CV
Comments 20 pages,6 figures,survey
机构 * Mila - Quebec AI Institute(魁北克AI研究院) ; Université de Montréal(蒙特利尔大学) ; McGill University(麦吉尔大学) ; Meta FAIR ; Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG
Comments Published at ICCV 2025
专题命中 视觉定位与Grounding :multimodal large language model(title);MLLM(abstract);分类 cs.CV
机构 * Logic Programming and Argumentation Group, TU Dresden, Germany(图灵编程与论证组,德累斯顿理工大学,德国) ; Knowledge-Based Systems Group, TU Dresden, Germany(知识系统组,德累斯顿理工大学,德国) ; DIMES - University of Calabria, Italy(迪梅斯-卡拉布里亚大学,意大利)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI
机构 * School of Automation Science and Engineering, South China University of Technology(自动化科学与工程学院,华南理工大学) ; School of Future Technology, South China University of Technology(未来技术学院,华南理工大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(title,abstract)
专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI
Comments arXiv admin note: text overlap with arXiv:2505.04410
机构 * stu.suda.edu.cn(苏州大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
机构 * Delft University of Technology(代尔夫特理工大学)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI
Comments Accepted for AIED 2025
机构 * University of Birmingham(伯明翰大学) ; Queen Mary University of London(伦敦大学女王学院) ; Aston University(阿斯顿大学) ; United Kingdom National Nuclear Laboratory Ltd.(英国国家核实验室有限公司)
专题命中 GUI与屏幕智能体 :vision-language model(title,abstract);VLM(abstract);分类 cs.AI
Comments Accepted at Human-Centered Robot Autonomy for Human-Robot Teams (HuRoboT) at IEEE RO-MAN 2025, Eindhoven, the Netherlands
专题命中 GUI与屏幕智能体 :grounding(abstract);multimodal large language model(abstract);分类 cs.CV
机构 * Zhejiang University(浙江大学) ; vivo AI Lab(vivo AI实验室)
专题命中 GUI与屏幕智能体 :VLM(abstract);分类 cs.AI
Comments Accepted by ACM MM 2025
机构 * South China University of Technology(华南理工大学) ; Pazhou Lab(Pazhou 实验室)
专题命中 GUI与屏幕智能体 :vision language model(abstract)
Comments ACM MM 2025
机构 * Togo AI Labs(Togo人工智能实验室) ; Vizuara AI Labs(Vizuara人工智能实验室)
专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI
机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) ; National Engineering Research Center for Visual Information and Applications(视觉信息与应用国家工程研究中心) ; Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院) ; Xi’an Jiaotong University(西安交通大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Osaka University(大阪大学)
专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV、cs.AI
Comments Accepted by IEEE Transactions on Information Forensics & Security
机构 * Institute of Theoretical and Applied Informatics, Polish Academy of Sciences(波兰科学院理论与应用信息学研究所) ; University of Lille, CNRS, Inria, Centrale Lille, UMR 9189 CRIStAL(里尔大学)
专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.LG
Comments 12 pages, 4 figures, 3 tables. Submitted for peer review
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(title,abstract);grounding(abstract)
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);分类 cs.CV
机构 * Information Materials and Intelligent Sensing Laboratory of Anhui Province(安徽省信息材料与智能感知实验室) ; Anhui Provincial Key Laboratory of Multimodal Cognitive Computation(安徽省多模态认知计算重点实验室) ; School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院)
专题命中 VLM训练与架构 :VLM(title);vision-language model(abstract);分类 cs.CV
机构 * ByteDance(字节跳动)
专题命中 VLM训练与架构 :vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
机构 * Netherlands Cancer Institute University of Amsterdam(荷兰癌症研究所 阿姆斯特丹大学)
专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
Comments 10 pages, 2 figures, 2 tables, submitted at MICCAI IMIMIC workshop
机构 * AI Chip Center for Emerging Smart Systems(新兴智能系统人工智能芯片中心) ; The Hong Kong University of Science(香港科学大学) ; Zhejiang University College of Computer Science(浙江大学计算机科学学院)
专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV
机构 * Mediterranean Agronomic Institute of Montpellier - CIHEAM-IAMM(地中海农业研究院-CIHEAM-IAMM) ; Inria(法国国家信息与自动化技术研究院) ; INRAE(法国国家农业研究咨询中心) ; Cirad(国际热带农业研究中心) ; UMR TETIS(TETIS联合研究单位) ; Univ. of Montpellier(蒙彼利埃大学)
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
Comments Accepted at WACV'25