Vision-language models for chest radiography do not always need the image
胸部X光片的视觉-语言模型并不总是需要图像
Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh
机构
*
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-埃尔朗根-纽伦堡大学模式识别实验室)
;
Department of Diagnostic and Interventional Radiology, TUM University Clinic, School of Medicine and Health, Klinikum rechts der Isar, Technical University of Munich(慕尼黑工业大学医学院与健康学院伊萨尔河右岸医院诊断与介入放射学系)
;
Lab for AI in Medicine, RWTH Aachen University(亚琛工业大学医学人工智能实验室)
;
Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen(亚琛工业大学医院诊断与介入放射学系)
Hallucination Detection and Correction in Medical VLMs via Counter-Evidence Verification
基于反事实证据验证的医学视觉语言模型幻觉检测与纠正
Nan Zhou, Ke Zou, Meng Liu, Linchao He, Jiaqi Zhu, Yi Zhang, Hu Chen, Huazhu Fu
机构
*
College of Computer Science, Sichuan University(四川大学计算机科学学院)
;
Yong Loo Lin School of Medicine, National University of Singapore(新加坡国立大学杨潞龄医学院)
;
Key Laboratory of Data Protection and Intelligent Management, Ministry of Education, Sichuan University(四川大学数据保护与智能管理教育部重点实验室)
;
National Key Laboratory of Autonomous Intelligent Unmanned Systems, Beijing Institute of Technology(北京理工大学自主智能无人系统国家重点实验室)
;
Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局高性能计算研究所)
CommentsInitial controlled diagnostic study on 23 natural drawing sets and three VLMs; broader model, building, repeated-inference, and human coverage is planned for a subsequent version
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
University of Freiburg(弗莱堡大学)
;
Hangzhou City University(杭州城市大学)
;
XGRIDS(XGRIDS公司)
;
Fudan University(复旦大学)
;
Hong Kong Baptist University(香港浸会大学)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving
UniDrive: 面向自动驾驶可解释风险理解的统一视觉-语言与定位框架
Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
机构
*
organization= Department of Earth Science \& Engineering, Imperial College London , city= London , postcode= SW7 2AZ , country= United Kingdom
;
organization= SpaceTimeLab, Department of Civil, Environmental
;
Geomatic Engineering, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom
;
organization= Department of Computing, The Hong Kong Polytechnic University , city= Hong Kong , country= China
;
organization= Trinity College, University of Oxford , city= Oxford , postcode= OX1 3BH , country= United Kingdom
;
organization= Department of Geography, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom
;
organization= Centre for Global Infrastructure Resilience, The Bartlett School of Sustainable Construction, University College London , city= London , postcode= WC1E 7HB , country= United Kingdom
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
ID-VTG: Image-Disambiguated Video Temporal Grounding
ID-VTG:基于图像消歧的视频时间定位
Minghang Zheng, Jingli Wei, Hongyi Yang, Yang Liu
机构
*
Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所)
;
State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation
RRS-10K:用于罕见遥感图像解释的多任务视觉语言模型基准测试
Yuqiao Lai, Jiancheng Qi, Fei Wang, Yuxin Liu, Kun Li, Ye Chen, Yan Gao, Yanyan Wei
机构
*
Laboratory of Intelligent Language Processing, National University of Defense Technology(国防科技大学智能语言处理实验室)
;
Hefei University of Technology(合肥工业大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
;
United Arab Emirates University(阿联酋大学)