UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning
UltraVR:面向证据推理的诊断性超分辨率图像VQA基准
Gexin Huang, Yanting Yang, Myeongkyun Kang, Beidi Zhao, Jun Zhou, Chen Zhou, Gang Wang, Zu-hua Gao, Xiaoxiao Li
机构
*
University of British Columbia(不列颠哥伦比亚大学)
;
Vector Institute(向量研究所)
;
BC Cancer Agency(不列颠哥伦比亚癌症中心)
;
The Hong Kong Polytechnic University(香港理工大学)
SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomical Reasoning in Radiology Vision-Language Models
SPARC-Rad:面向放射学视觉语言模型的空间与解剖推理多模态基准数据集及评估流程
Satvik Tripathi, Mustafa Ege Seker, Kristian Quevada, Ebubechukwu D Enwerem, Pratham Khandelwal, Emine Meltem, Bera Koca, Shahriar Faghani, Jacinta Arnold, Dania Daye, Tessa S. Cook
Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering
用于通用文档视觉问答的领域适应视觉语言模型的比较研究
Miguel Lopez-Duran, Elena Marrero, Julian Fierrez, Marta Robledo-Moreno, Ruben Vera-Rodriguez, Daniel DeAlcala, Aythami Morales, Ruben Tolosana, Oscar Delgado, Alvaro Ortigosa, Javier Ortega-Garcia
机构
*
Universidad Autónoma de Madrid (UAM)(马德里自治大学)
;
BiometricsAI(生物识别人工智能)
机构
*
Computer Science Department, Morgan State University(莫尔甘州大学计算机科学系)
;
International Organization for Migration (IOM)(国际移民组织)
;
Electrical & Computer Engineering Department, Morgan State University(莫尔甘州大学电气与计算机工程系)
UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation
UCSF-PDGM-VQA: 用于脑肿瘤MRI解读的视觉问答数据集
Shiv Ghosh, Junayd Lateef, Chih-Hua Liu, Yannan Yu, Andreas M. Rauschecker, Madhumita Sushil
机构
*
Fung Institute for Engineering Leadership(工程领导力基金会)
;
University of California, Berkeley(加州大学伯克利分校)
;
Department of Radiology(放射科)
;
University of California, San Francisco(加州大学旧金山分校)
;
Division of Clinical Informatics and Digital Transformation(临床信息学与数字转型部)
;
Department of Neurological Surgery(神经外科部)
Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
迈向具有解释链预测的自解释文档视觉问答
Kjetil Indrehus, Adrian Duric, Changkyu Choi, Ali Ramezani-Kebrya
机构
*
Department of Informatics, University of Oslo(奥斯陆大学信息学院)
;
Integreat – Norwegian Centre for Knowledge-driven Machine Learning(挪威知识驱动机器学习中心)
;
TRUST – The Norwegian Centre for Trustworthy AI(挪威可信人工智能中心)
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
专题命中
视觉问答
:visual question answering(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
机构
*
Interdisciplinary Program in Artificial Intelligence(人工智能交叉学科项目)
;
Department of Intelligence and Information(智能与信息系)
;
Samsung Advanced Institute of Technology(三星先进技术研究院)
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
JaWildText:面向日文场景文本理解的视觉-语言模型基准测试
Koki Maeda, Naoaki Okazaki
机构
*
Institute of Science Tokyo, Tokyo, Japan(东京科学大学,日本东京)
;
Research and Development Center for Large Language Models, National Institute of Informatics, Tokyo, Japan(国立信息学研究所大型语言模型研发中心,日本东京)
机构
*
University of California, San Diego(加利福尼亚大学圣迭戈分校)
;
University of California, Riverside(加利福尼亚大学河滨分校)
;
University of Southern California(南加利福尼亚大学)
CommentsWe request withdrawal of this paper because one of the listed institutional affiliations was included without proper authorization. This issue cannot be resolved through a simple revision, and we therefore request withdrawal to prevent dissemination of incorrect or unauthorized affiliation information
RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
RadImageNet-VQA:一个大规模的CT和MRI数据集用于放射学视觉问答
Léo Butsanets, Charles Corbière, Julien Khlaut, Pierre Manceron, Corentin Dancette
机构
*
Raidium
;
Université de Paris Cité, Hôpital Européen Georges Pompidou, AP-HP(巴黎西岱大学,乔治·蓬皮杜欧洲医院,AP-HP)
;
Department of Vascular and Oncological Interventional Radiology, INSERM(血管与肿瘤介入放射学系,法国国家健康与医学研究院)
Rethinking Token Reduction for Large Vision-Language Models
重新思考大型视觉-语言模型中的标记缩减
Yi Wang, Haofei Zhang, Qihan Huang, Anda Cao, Gongfan Fang, Wei Wang, Xuan Jin, Jie Song, Mingli Song, Xinchao Wang
机构
*
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江大学区块链与数据安全国家重点实验室)
;
Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术区(滨江)区块链与数据安全研究院)
;
School of Software Technology, Zhejiang University(浙江大学软件学院)
;
National University of Singapore(新加坡国立大学)
;
Alibaba Group(阿里巴巴集团)