arXivDaily arXiv每日学术速递 周一至周五更新

科学与医疗

医学 AI

医学智能、临床 AI、医学影像、病理、诊断和医疗健康大模型。

共收录 975 信号源:cs.CV, cs.LG, q-bio, eess.IV, eess.SP

1. 医疗多模态 975 篇

2604.26288 2026-05-01 cs.CV cs.AI 57%

CheXthought: A global multimodal dataset of clinical chain-of-thought reasoning and visual attention for chest X-ray interpretation

CheXthought:一个包含临床推理链和视觉注意力的全球多模态数据集,用于胸部X光解读

Sonali Sharma, Jin Long, George Shih, Sarah Eid, Christian Bluethgen, Francine L. Jacobson, Emily B. Tsai, Global Radiology Consortium, Ahmed M. Alaa, Curtis P. Langlotz

机构 * Center for Artificial Intelligence in Medicine and Imaging(医学与影像人工智能中心) Department of Radiology(放射科) Department of Pediatrics(儿科) Department of Radiology, Weill Cornell Medicine(韦氏 Cornell 医学院放射科) Department of Radiology, Brigham and Women’s Hospital(布里奇妇产科医院放射科) University of California, Berkeley(加州大学伯克利分校) University of California, San Francisco(加州大学旧金山分校) Department of Biomedical Data Science(生物医学数据科学部)

专题命中 医疗多模态 :pathology(abstract);分类 cs.CV

AI总结 CheXthought数据集通过501名放射科医生的50312例多读X光片,提供了103592条推理链和6609082条同步视觉注意力标注,展示了临床推理模式,并在事实准确性、空间定位、病理分类、视觉忠实度、时间推理和不确定性沟通等方面提升了视觉语言模型性能。

Comments 51 pages, 7 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25720 2026-04-29 cs.CV cs.CL 57%

Toward Multimodal Conversational AI for Age-Related Macular Degeneration

迈向年龄相关性黄斑变性的多模态对话式人工智能

Ran Gu, Benjamin Hou, Mélanie Hébert, Asmita Indurkar, Yifan Yang, Emily Y. Chew, Tiarnán D. L. Keenan, Zhiyong Lu

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 本文提出OcularChat,一种基于多模态大语言模型的系统,通过模拟患者与医生对话,利用视网膜彩色照相进行黄斑变性诊断,展现出优于现有模型的分类性能和临床解释能力。

Comments 38 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20306 2026-04-23 cs.CV cs.AI 57%

Dual Causal Inference: Integrating Backdoor Adjustment and Instrumental Variable Learning for Medical VQA

双重因果推断:整合后门调整与工具变量学习用于医疗视觉问答

Zibo Xu, Qiang Li, Ke Lu, Jin Wang, Weizhi Nie, Yuting Su

机构 * School of Microelectronics, Tianjin University(天津大学微电子学院) Department of Orthopedics, Affiliated Kunshan Hospital of Jiangsu University(江苏大学附属昆山医院骨科部) Department of Clinical Laboratory, The Third Central Hospital of Tianjin(天津第三中心医院临床实验室) School of Electrical and Information Engineering, Tianjin University(天津大学电气与信息工程学院)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 本文提出双重因果推断框架,通过整合后门调整和工具变量学习,解决医疗视觉问答中多模态数据中的内在偏见问题,提升模型的诊断推理可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03740 2026-04-23 cs.CV cs.CL 57%

CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values

CLIP-SVD:通过奇异值实现高效且可解释的视觉-语言适应

Taha Koleilat, Hassan Rivaz, Yiming Xiao

机构 * Department of Electrical & Computer Engineering, Concordia University(康科迪亚大学电气与计算机工程系) Department of Computer Science & Software Engineering, Concordia University(康科迪亚大学计算机科学与软件工程系)

专题命中 医疗多模态 :biomedical(abstract);分类 cs.CV

AI总结 CLIP-SVD通过奇异值微调方法,以0.04%的参数量实现高效视觉-语言模型适应,提升适应性能并保留泛化能力,在11个自然和10个生物医学数据集上取得最佳分类结果。

Comments TMLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18757 2026-04-22 cs.CV cs.AI 57%

REVEAL: Multimodal Vision-Language Alignment of Retinal Morphometry and Clinical Risks for Incident AD and Dementia Prediction

REVEAL:多模态视图-语言对齐的视网膜形态学与临床风险预测

Seowung Leem, Lin Gu, Chenyu You, Kuang Gong, Ruogu Fang

机构 * J. Crayton Pruitt Family Department of Biomedical Engineering, University of Florida(佛罗里达大学J. Crayton Pruitt家族生物医学工程系) Research Institute of Electrical Communication, Tohoku University(东北大学电气通信研究所) Department of Applied Mathematics & Statistics, Stony Brook University(石溪大学应用数学与统计学系) Department of Computer Science, Stony Brook University(石溪大学计算机科学系)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 REVEAL通过多模态对齐视网膜形态学与临床风险因素,提前8年预测阿尔茨海默病和痴呆症,优于现有模型。

Comments Accepted for publication a MIDL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17570 2026-04-21 cs.CV cs.AI 57%

PBSBench: A Multi-Level Vision-Language Framework and Benchmark for Hematopathology Whole Slide Image Interpretation

PBSBench:一种多级视觉-语言框架和血液病理学全滑片图像解释基准

Yuanlong Wang, Weichi Chen, Adrian Rajab, Wenfang Liu, Yulan Jin, Andrew Srisuwananukorn, Ping Zhang

机构 * The Ohio State University(俄亥俄州立大学) The Ohio State University Wexner Medical Center(俄亥俄州立大学韦克斯纳医学中心)

专题命中 医疗多模态 :pathology(abstract);分类 cs.CV

AI总结 本文提出PBSBench,一种针对血液病理学全滑片图像解释的多级视觉-语言框架和基准,通过构建PBSInstr数据集和PBS-VL模型,提升了对血涂片图像的理解能力。

Comments 19 pages, 12 figures, Accepted by CVPR Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22278 2026-04-20 cs.CV 57%

FETAL-GAUGE: A Benchmark for Assessing Vision-Language Models in Fetal Ultrasound

FETAL-GAUGE:用于评估胎儿超声检查中视觉-语言模型的基准

Hussain Alasmawi, Numan Saeed, Mohammad Yaqub

机构 * Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 本文提出Fetal-Gauge基准,用于评估视觉-语言模型在胎儿超声检查中的性能,包含42000张图像和93000个问题-答案对,揭示当前模型在临床任务中的性能差距,强调需领域适应的架构和训练方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14866 2026-04-17 cs.CV cs.AI 57%

MetaDent: Labeling Clinical Images for Vision-Language Models in Dentistry

MetaDent: 临床图像标注用于牙科领域视觉-语言模型

Meng-Xun Li, Wen-Hui Deng, Zhi-Xing Wu, Chun-Xiao Jin, Jia-Min Wu, Yue Han, James Kit Hon Tsoi, Gui-Song Xia, Cui Huang

机构 * School of Computer Science, Wuhan University, Wuhan, Hubei, China(计算机科学学院,武汉大学,武汉,湖北,中国)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 本文提出MetaDent,通过大规模牙科图像数据集、半结构化标注框架和综合基准测试,解决牙科内窥镜图像分析中缺乏细粒度标注的问题,验证了VLMs在临床图像理解中的性能局限。

Comments Project website: https://menxli.github.io/metadent

Journal ref Journal of Dental Research, p.00220345261424242 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14656 2026-04-17 cs.AI cs.CL cs.CV 57%

Rethinking Patient Education as Multi-turn Multi-modal Interaction

重新思考患者教育作为多轮多模态交互

Zonghai Yao, Zhipeng Tang, Chengtao Lin, Xiong Luo, Benlu Wang, Juncheng Huang, Chin Siang Ong, Hong Yu

机构 * VA Bedford Health Care(VA贝德福德医疗中心) UMass Amherst(马萨诸塞大学阿默斯特分校) UMass Lowell(马萨诸塞大学洛厄尔分校) Yale University(耶鲁大学) National University of Singapore(新加坡国立大学) Yale School of Medicine(耶鲁医学院)

专题命中 医疗多模态 :radiology(abstract);分类 cs.CV

AI总结 本文提出MedImageEdu基准,通过多轮多模态交互提升患者教育效果,评估咨询过程和最终响应质量,发现多模态模型在视觉 grounding、安全性和情绪互动方面存在不足。

Comments Equal contribution for the first two authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21652 2026-04-15 eess.IV cs.AI physics.med-ph 57%

Enabling Ultra-Fast Cardiovascular Imaging Across Heterogeneous Clinical Environments with A Generalist Foundation Model and Multimodal Database

通过通用基础模型和多模态数据库实现跨异构临床环境的超快速心血管成像

Zi Wang, Mingkai Huang, Zhang Shi, Hongjie Hu, Lan Lan, Hui Zhang, Yan Li, Xi Hu, Qing Lu, Zongming Zhu, Qiong Yao, Yuxiang Dai, Fanwen Wang, Yinzhe Wu, Jun Lyu, Qianqian Gao, Guangming Xu, Zhenxuan Zhang, Haosen Zhang, Qing Li, Guangming Wang, Tianxing He, Lizhen Lan, Siyue Li, Le Xue, Mengting Sun, Yuntong Lyu, Junpu Hu, Jiayu Zhu, Rizwan Ahmad, Zhengyu Bu, Xianling Qian, Guanke Cai, Ruiyu Cao, Weirui Cai, Chang Xu, Yuyang Ren, Feidan Yu, Siying Ma, Ziqiang Xu, Xinran Chen, Sha Hua, Daniel Kim, Yajing Zhang, Chen Ouyang, Wenjia Bai, Jing Qin, Yucheng Yang, Daniel Rueckert, He Wang, Qian Tao, Claudia Prieto, Michael Markl, Alistair Young, Lianming Wu, Shuo Wang, Chen Qin, Mengsu Zeng, Xihong Hu, Haibo Xu, Xiaobo Qu, Hao Li, Guang Yang, Chengyan Wang

专题命中 医疗多模态 :diagnosis(abstract);分类 eess.IV

AI总结 本文提出CardioMM模型,通过统一语义理解和物理约束数据一致性,实现跨不同扫描仪、协议和患者情况的稳健重建,验证了24倍加速仍能保持关键心脏表型和诊断图像质量。

Comments Github: https://github.com/wangziblake/CardioMM_MMCMR-427K

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09364 2026-04-14 cs.CV cs.CL 57%

Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts

仲裁失败,而非感知失明:视觉-语言模型如何解决视觉-语言冲突

Farhad Nooralahzadeh, Omid Rohanian, Yi Zhang, Jonathan Fürst, Kurt Stockinger

机构 * Institute of Computer Science, Zurich University of Applied Sciences(苏黎世应用科技大学计算机科学研究所) University of Oxford(牛津大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 研究探讨视觉-语言模型在视觉与语言冲突中的仲裁机制,发现编码与基础的脱节,通过多模态仲裁交叉分析揭示视觉属性在早期层可线性解码,最终层logit与基础结果相关性达0.847,表明需针对性干预提升视觉基础能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09841 2026-04-14 cs.CV cs.AI 57%

Is There Knowledge Left to Extract? Evidence of Fragility in Medically Fine-Tuned Vision-Language Models

还有可以提取的知识吗?医学微调视觉语言模型中的脆弱性证据

Oliver McLaughlin, Daniel Shubin, Carsten Eickhoff, Ritambhara Singh, William Rudman, Michal Golovanevsky

机构 * Brown University(布朗大学) University of Washington(华盛顿大学) University of Tübingen(蒂宾根大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 研究评估了四个医学微调的视觉语言模型在四个医学影像任务中的表现,发现任务难度增加时性能下降至接近随机水平,表明临床推理有限,医学微调无明显优势,模型对提示词敏感。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07814 2026-04-10 cs.CV 57%

AgriChain Visually Grounded Expert Verified Reasoning for Interpretable Agricultural Vision Language Models

AgriChain:基于视觉的专家验证推理用于可解释的农业视觉语言模型

Hazza Mahmood, Yongqiang Yu, Rao Anwer

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 本文提出AgriChain数据集,通过专家验证的推理链提升农业视觉语言模型的准确性和可解释性,实验表明其在植物病害诊断中优于多个基线模型。

Comments 9 pages

Journal ref LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04614 2026-04-09 cs.LG cs.AI 57%

A Clinical Point Cloud Paradigm for In-Hospital Mortality Prediction from Multi-Level Incomplete Multimodal EHRs

一种用于院内死亡预测的临床点云范式:从多级不完整多模态电子健康记录

Bohao Li, Tao Zou, Junchen Ye, Yan Gong, Bowen Du

机构 * Beihang University(北京航空航天大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.LG

AI总结 本文提出HealthPoint范式,通过统一的4D空间表示多级不完整多模态EHRs,引入低秩关系注意力机制和分层交互策略,实现灵活的事件级交互与细粒度自监督,提升模态恢复和未标记数据利用效率,实验表明其在风险预测中性能优异。

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18297 2026-04-08 cs.CV 57%

Image-to-Text for Medical Reports Using Adaptive Co-Attention and Triple-LSTM Module

利用自适应共注意和三LSTM模块的图像到文本生成医学报告

Yishen Liu

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 本文提出CA-TriNet模型,结合Transformer和多LSTM网络,通过自适应权重运算和三LSTM模块提升医学图像与文本生成的准确性与多样性。

Comments arXiv admin note: This submission has been withdrawn by arXiv administrators due to incorrect authorship. Author list truncated

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05748 2026-04-08 cs.CV 57%

SVC 2026: the Second Multimodal Deception Detection Challenge and the First Domain Generalized Remote Physiological Measurement Challenge

SVC 2026: 第二届多模态欺骗检测挑战赛及首个领域通用远程生理测量挑战

Dongliang Zhu, Zhiyi Niu, Bo Zhao, Jiajian Huang, Shuo Ye, Xun Lin, Hui Ma, Taorui Wang, Jiayu Zhang, Chunmei Zhu, Junzhe Cao, Yingjie Ma, Rencheng Song, Albert Clapés, Sergio Escalera, Dan Guo, Zitong Yu

机构 * Wuhan University(武汉大学) Great Bay University(大湾区大学) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学) Sun Yat-sen University(中山大学) Hefei University of Technology(合肥工业大学) University of Barcelona(巴塞罗那大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 本文提出SVC 2026挑战赛,旨在通过多模态欺骗检测和远程脉搏波测速估计任务,推动对细微视觉信号的鲁棒表示学习研究。

Comments Accepted by the SVC workshop @ CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05831 2026-04-08 cs.LG cs.AI 57%

HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding

HeartcareGPT:一种统一的多模态ECG套件,用于双信号-图像建模与理解

Yihan Xie, Sijing Li, Tianwei Lin, Zhuonan Wang, Chenglin Yang, Yu Zhong, Wenjie Yan, Wenqiao Zhang, Xiaogang Guo, Jun Xiao, Yueting Zhuang, Beng Chin Ooi

机构 * Zhejiang University(浙江大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.LG

AI总结 本文提出HeartcareGPT,通过Dual Stream Projection Alignment机制,实现ECG信号与图像的统一建模,提升多视角ECG理解能力,并建立医学多模态大语言模型向生理信号领域扩展的方法论基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04133 2026-04-07 cs.CV cs.AI 57%

Learning Robust Visual Features in Computed Tomography Enables Efficient Transfer Learning for Clinical Tasks

在计算机断层扫描中学习鲁棒的视觉特征以实现临床任务的高效迁移学习

Rubén Moreno-Aguado, Alba Magallón, Victor Moreno, Yingying Fang, Guang Yang

机构 * Bioengineering Department and Imperial-X, Imperial College London(帝国理工学院生物工程系与Imperial-X) Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系) Oncology Data Analytics Program, Catalan Institute of Oncology(加泰罗尼亚肿瘤研究所肿瘤数据分析项目) Colorectal Cancer Group, ONCOBELL Program, Institut d’Investigació Biomèdica de Bellvitge(贝尔维特奇生物医学研究所ONCOBELL项目结直肠癌研究组) Consortium for Biomedical Research in Epidemiology and Public Health(流行病学与公共卫生生物医学研究联盟) Department of Clinical Sciences, Faculty of Medicine and Health Sciences, Universitat de Barcelona(巴塞罗那大学医学与健康科学学院临床科学系) Institute of Complex Systems, University of Barcelona(巴塞罗那大学复杂系统研究所) National Heart and Lung Institute, Imperial College London(帝国理工学院国家心肺研究所) Cardiovascular Research Centre, Royal Brompton Hospital(皇家布朗普顿医院心血管研究中心) School of Biomedical Engineering & Imaging Sciences, King’s College London(伦敦国王学院生物医学工程与影像科学学院)

专题命中 医疗多模态 :CT(abstract);分类 cs.CV

AI总结 本文提出VoxelFM,一种基于自蒸馏的3D CT基础模型,通过学习语义丰富的特征,实现了在多种临床任务中无需微调即可高效迁移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17514 2026-04-07 cs.CV 57%

EI: Early Intervention for Multimodal Imaging based Disease Recognition

EI:基于多模态影像的疾病识别早期干预

Qijie Wei, Hailan Lin, Xirong Li

机构 * Renmin University of China(中国人民大学) Beijing Key Laboratory for Intelligent Diagnosis of Fundus Diseases and Drug-Device R&D and Translation(眼底疾病智能诊断与药械研发转化北京市重点实验室)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 本文提出EI框架,通过早期干预提升多模态影像疾病识别性能,引入MoR方法优化Vision Foundation Models适应,验证了在三个公开数据集上的有效性。

Comments Accepted to CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02748 2026-04-06 cs.CV 57%

Visual Instruction-Finetuned Language Model for Versatile Brain MR Image Tasks

面向多样化脑部MRI任务的视觉指令微调语言模型

Jonghun Kim, Sinyoung Ra, Hyunjin Park

机构 * Sungkyunkwan University(成均馆大学) Department of Electrical and Computer Engineering(电气与计算机工程系) Department of Artificial Intelligence(人工智能系)

专题命中 医疗多模态 :MRI(abstract);分类 cs.CV

AI总结 本文提出LLaBIT模型,通过视觉指令微调提升语言模型在脑部MRI任务中的表现,通过特征图重用和文本数据生成提升性能,验证了其在报告生成、视觉问答、图像分割和图像翻译中的有效性。

Comments ICPR 2026 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01310 2026-04-03 cs.CV 57%

Sparse Spectral LoRA: Routed Experts for Medical VLMs

稀疏光谱LoRA:用于医学VLMs的路由专家

Omid Nejati Manzari, Hojat Asgariandehkordi, Taha Koleilat, Yiming Xiao, Hassan Rivaz

机构 * Concordia University(康考迪亚大学)

专题命中 医疗多模态 :radiology(abstract);分类 cs.CV

AI总结 本文提出MedQwen,一种参数高效的医学VLM,结合了光谱路由的专家混合(MoE)与理论支持的缩放规则,通过非重叠SVD段初始化专家,减少序列遗忘并提升医疗图像任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26726 2026-03-31 cs.CV cs.AI 57%

A Multimodal Deep Learning Framework for Edema Classification Using HCT and Clinical Data

一种结合头颅CT和临床数据的多模态深度学习框架用于水肿分类

Aram Ansary Ogholbake, Hannah Choi, Spencer Brandenburg, Alyssa Antuna, Zahraa Al-Sharshahi, Makayla Cox, Haseeb Ahmed, Jacqueline Frank, Nathan Millson, Luke Bauerle, Jessica Lee, David Dornbos, Qiang Cheng

机构 * University of Kentucky(肯塔基大学) Department of Computer Science, University of Kentucky(肯塔基大学计算机科学系) Department of Neurological Surgery, University of Kentucky(肯塔基大学神经外科学系) University of Kentucky College of Medicine(肯塔基大学医学院) Department of Neurology, University of Kentucky(肯塔基大学神经学系)

专题命中 医疗多模态 :CT(abstract);分类 cs.CV

AI总结 本文提出AttentionMixer框架,通过融合头颅CT和临床数据实现脑水肿检测,采用自监督Vision Transformer Autoencoder编码CT数据,并利用交叉注意力模块整合临床信息,最终通过MLP-Mixer提升分类性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26008 2026-03-30 cs.CV cs.AI 57%

FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants

FairLLaVA: 大规模视觉-语言助手中的公平性感知参数高效微调

Mahesh Bhosale, Abdul Wasi, Shantam Srivastava, Shifa Latif, Tianyu Luan, Mingchen Gao, David Doermann, Xuan Gong

机构 * University at Buffalo(布法罗大学) University of Kashmir(克什米尔大学) Accenture(埃森哲) Harvard Medical School(哈佛医学院)

专题命中 医疗多模态 :radiology(abstract);分类 cs.CV

AI总结 FairLLaVA通过减少目标属性间的互信息,实现视觉指令微调中的公平性改进,提升医疗影像生成的公平性和自然语言生成质量。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23953 2026-03-27 cs.CV cs.ET 57%

VOLMO: Versatile and Open Large Models for Ophthalmology

VOLMO:面向眼科学的多功能和开放型大模型

Zhenyue Qin, Younjoon Chung, Elijah Lee, Wanyue Feng, Xuguang Ai, Serina Applebaum, Minjie Zou, Yang Liu, Pan Xiao, Mac Singer, Amisha Dave, Aidan Gilson, Tiarnan D. L. Keenan, Emily Y. Chew, Zhiyong Lu, Yih-Chung Tham, Ron Adelman, Luciano V. Del Priore, Qingyu Chen

机构 * Department of Biomedical Informatics & Data Science, Yale University(耶鲁大学生物医学信息学与数据科学系) Ray and Stephanie Lane Computational Biology Department, Carnegie Mellon University(卡内基梅隆大学雷和斯蒂芬妮·兰德计算生物学系) Yong Loo Lin School of Medicine, National University of Singapore(新加坡国立大学杨洛林医学院) Department of Radiology, Washington University in Saint Louis(圣路易斯华盛顿大学放射科) National Eye Institute, National Institutes of Health(国家卫生研究院眼科研究所) National Library of Medicine, National Institutes of Health(国家卫生研究院国家医学图书馆)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 VOLMO提出了一种通用框架,用于开发专门的眼科多模态大语言模型,通过预训练、微调和临床推理三个阶段,提升了眼科疾病筛查和诊断的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13119 2026-03-26 cs.CV cs.AI 57%

Geometry-Guided Camera Motion Understanding in VideoLLMs

基于几何的视频LLMs中相机运动理解

Haoan Feng, Sri Harsha Musunuri, Guan-Ming Su

机构 * University of Maryland, College Park(马里兰大学College Park分校) Dolby Laboratories Inc.(杜比实验室)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 本文提出通过基准测试、诊断和注入框架,解决视频LLMs中相机运动表示不足的问题,通过CameraMotionDataset和CameraMotionVQA基准,提升模型对相机运动的识别能力。

Comments 10 pages, 7 figures, supplementary included CVPR2026 PVUW

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21597 2026-03-25 cs.AI cs.CV 57%

Cerebra: A Multidisciplinary AI Board for Multimodal Dementia Characterization and Risk Assessment

Cerebra:一个多学科AI平台用于多模态痴呆症特征化和风险评估

Sheng Liu, Long Chen, Zeyun Zhao, Qinglin Gou, Qingyue Wei, Arjun Masurkar, Kevin M. Spiegler, Philip Kuball, Stefania C. Bray, Megan Bernath, Deanna R. Willis, Jiang Bian, Lei Xing, Eric Topol, Kyunghyun Cho, Yu Huang, Ruogu Fang, Narges Razavian, James Zou

机构 * Department of Biomedical Data Science, Stanford University(斯坦福大学生物医学数据科学系) Department of Radiation Oncology, Stanford University(斯坦福大学放射肿瘤科) Center for Data Science, New York University(纽约大学数据科学中心) J. Crayton Pruitt Family Department of Biomedical Engineering, University of Florida(佛罗里达大学杰·克雷顿·普瑞特家庭生物医学工程系) Department of Biomedical Engineering and Informatics, Indiana University Indianapolis(印第安纳大学印第安纳波利斯分校生物医学工程与信息学系) Department of Neurology, NYU Grossman School of Medicine(纽约大学格罗斯曼医学院神经科) Department of Neuroscience, NYU Grossman School of Medicine(纽约大学格罗斯曼医学院神经科学系) UF Health Family Medicine – Haile Plantation(佛罗里达大学健康家庭医学-海莱种植园) Department of Family Medicine, Indiana University School of Medicine(印第安纳大学医学院家庭医学系) Department of Biostatistics and Health Data Science, Indiana University School of Medicine(印第安纳大学医学院生物统计学与健康数据科学系) Scripps Research Translational Institute(斯克里普斯研究转化研究所) Courant Institute, New York University(纽约大学柯朗研究所) Center for Biomedical Informatics, Regenstrief Institute(再生斯蒂夫研究所生物医学信息中心) Center for Cognitive Aging and Memory, McKnight Brain Institute, University of Florida(佛罗里达大学麦金托神经研究所认知衰老与记忆中心) Population Health Department, NYU Langone Health(纽约大学兰戈恩健康人口健康部) Radiology Department, NYU Langone Health(纽约大学兰戈恩健康放射科)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 Cerebra通过多智能体协调处理电子病历、临床笔记和医学影像,提供可视化分析和对话界面,提升临床决策支持的可解释性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21700 2026-03-24 cs.CV 57%

PPGL-Swarm: Integrated Multimodal Risk Stratification and Hereditary Syndrome Detection in Pheochromocytoma and Paraganglioma

PPGL-Swarm:集成多模态风险分层与遗传综合征检测的pheochromocytoma和paraganglioma诊断系统

Zelin Liu, Xiangfu Yu, Jie Huang, Ge Wang, Yizhe Yuan, Zhenyu Yi, Jing Xie, Haotian Jiang, Lichi Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Ruijin Hospital(瑞金医院) ShanghaiTech University(上海科技大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 本文提出PPGL-Swarm系统,通过多模态证据整合实现全面诊断报告,包括自动化GAPP评分、基因风险警报及可追溯的推理路径,解决传统诊断方法的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12469 2026-03-24 cs.CV 57%

Unleashing Video Language Models for Fine-grained HRCT Report Generation

释放视频语言模型以生成细粒度HRCT报告

Yingying Fang, Huichi Zhou, KinHei Lee, Yijia Wang, Zhenxuan Zhang, Jiahao Huang, Guang Yang

机构 * Bioengineering Department, Imperial College London, London, UK School of Biomedical Engineering \& lmaging Sciences, King's College London, London,UK

专题命中 医疗多模态 :CT(abstract);分类 cs.CV

AI总结 本文提出AbSteering框架,通过异常中心方案和直接偏好优化目标,提升视频语言模型在HRCT报告生成中的精度与细粒度区分能力,优于现有领域特定CT基础模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16664 2026-03-18 cs.CV cs.AI 57%

Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation

Kestrel: 为降低LVLM幻觉而引入自反思

Jiawei Mao, Hardy Chen, Haoqin Tu, Yuhan Wang, Letian Zhang, Zeyu Zheng, Huaxiu Yao, Zirui Wang, Cihang Xie, Yuyin Zhou

机构 * UC Santa Cruz(加州大学圣克ruz分校) UC Berkeley(加州大学伯克利分校) UNC-Chapel Hill(北卡罗来纳大学教堂山分校) Apple(苹果公司)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 Kestrel提出一种无需训练的框架,通过显式视觉 grounding 与证据验证自反思机制减少LVLM幻觉,实验显示在POPE和MME-Hallucination基准上性能提升,同时提供透明的验证轨迹。

Comments 16 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16372 2026-03-18 cs.CV 57%

InViC: Intent-aware Visual Cues for Medical Visual Question Answering

InViC:面向医疗视觉问答的意图感知视觉线索

Zhisong Wang, Ziyang Chen, Zanting Ye, Hongze Zhu, Yefeng Zheng, Yong Xia

机构 * National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, School of Computer Science and Engineering, Northwestern Polytechnical University, Xi’an 710072, China(集成空天地海大数据应用技术国家工程实验室,计算机科学与工程学院,西北工业大学,西安710072,中国) Westlake University(西湖大学) Southern Medical University(南方医科大学)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 本文提出InViC框架,通过引入Cue Tokens Extraction模块和两阶段微调策略,提升医疗视觉问答中意图对齐的视觉证据利用,验证了瓶颈化训练在提高可信度上的有效性。

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏