arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26403 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7473 篇

2606.30247 2026-06-30 cs.CL 90%

Grounding LLM Reasoning under Incomplete Graph Evidence

在不完全图证据下 grounding LLM 推理

Jiaqi Li, Fanghui Song

机构 * Tianjin Normal University, College of Computer and Information Engineering(天津师范大学计算机与信息工程学院) Harbin Institute of Technology, School of Mathematics(哈尔滨工业大学数学学院)

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 本文提出在不完全知识图谱证据下,通过KL正则化变形LLM先验实现软grounding,并给出稳定性界限,适用于GraphRAG、KGQA等场景。

Comments A theoretical perspective about Grounding LLM Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05978 2026-07-08 cs.CV cs.AI 新提交 90%

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention

提出并关注:通过多令牌局部注意力实现无训练的多模态大语言模型接地置信度

Daniel Shalam, Emanuel Ben Baruch, Avi Ben Cohen, Tal Remez

机构 * Amazon(亚马逊)

专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 研究多模态大语言模型接地置信度问题,提出无训练的多令牌局部注意力(MTLA)方法,通过在声称区域内求和并跨预测令牌聚合恢复更强接地信号,在多模型家族和模态中提升了幻觉AUROC及零样本检测AP。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28401 2026-06-30 cs.CV cs.LG 90%

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs

视觉驱动的偏好合成用于缓解VLM中的幻觉

Yunhun Nam, Jongheon Jeong

专题命中 视觉定位与Grounding :VLM(title_cn,summary_cn);vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.LG

AI总结 提出ViPSy框架,通过视觉线索构建策略对齐且视觉 grounded 的偏好数据,显著降低VLM幻觉率,在AMBER和Object HalBench上分别降低35.7%和24.5%,并提升通用视觉基准性能。

Comments 29 pages; Code is available at https://github.com/yunpal/ViPSy

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31196 2026-06-01 cs.CV cs.AI cs.CL cs.RO 90%

Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration

探索视觉-语言模型中的碰撞接地以实现安全的人机协作

Jun Wang, Xiaohao Xu, Xiaonan Huang

机构 * University of Michigan, Ann Arbor(密歇根大学,安娜堡)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 针对安全人机协作,提出碰撞接地概念及物理基准TouchSafeBench,评估视觉-语言模型在分类当前安全状态和预警即将碰撞任务中的表现,发现现有模型不可靠,视觉流畅性不等于物理责任性。

Comments 31 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22902 2026-05-25 cs.LG cs.AI cs.CL 90%

Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models

Transcoders 追踪视觉语言模型中的视觉基础与幻觉

Dimitrios Damianos, Leon Voukoutis, Georgios Skyrianos, Vassilis Katsouros, Georgios Paraskevopoulos

机构 * Institute of Language and Speech Processing(语言与语音处理研究所) Athena Research Center(雅典研究中心)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract_cn);分类 cs.AI、cs.LG

AI总结 采用基于Transcoders的功能中心框架分解视觉语言模型的计算路径,揭示视觉输入如何影响文本生成,并通过反事实分析和图结构特征预测幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00528 2026-04-03 cs.CV cs.AI 90%

Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding

思考、行动、构建:一种基于视觉语言模型的代理框架用于零样本3D视觉定位

Haibo Wang, Zihao Lin, Zhiyang Xu, Lifu Huang

机构 * University of California, Davis(加州大学戴维斯分校) Virginia Tech(弗吉尼亚理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision language model(title);vision-language model(abstract);VLM(abstract)

AI总结 本文提出TAB框架,通过2D到3D重建范式直接处理原始RGB-D流,利用2D VLMs解析复杂空间语义并结合多视图几何构建3D结构,克服传统方法依赖预处理点云的局限,实验证明其在ScanRefer和Nr3D上优于零样本方法和全监督基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01914 2026-03-24 cs.CV cs.AI cs.CL 90%

HPE-CogVLM: Advancing Vision Language Models with a Head Pose Grounding Task

HPE-CogVLM: 通过头部姿态接地任务提升视觉语言模型

Yu Tian, Tianqi Shao, Tsukasa Demizu, Xuyang Wu, Hsin-Tai Wu

机构 * Docomo Innovations, Inc.(Docomo创新公司) Department of Computer Science and Engineering, Santa Clara University(圣克拉拉大学计算机科学与工程系)

专题命中 视觉定位与Grounding :vision language model(title,abstract);grounding(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

AI总结 本文提出HPE-CogVLM框架,利用VLM的物体检测能力提升头部姿态估计精度,通过改进的LoRA层合并方法,有效解决融合任务中的响应格式问题,实现优于现有方法的性能。

Comments Accepted by IEEE Transactions on Circuits and Systems for Video Technology (TCSVT), 2026. This version includes major updates in methodology and experiments. The final version is available at IEEE Xplore

Journal ref IEEE Transactions on Circuits and Systems for Video Technology, Early Access, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12633 2026-03-18 cs.CV cs.AI 90%

DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model

DiG:通过差异 grounding 提升多模态大语言模型的细粒度感知

Zhou Tao, Shida Wang, Yongxiang Hua, Haoyu Cao, Linli Xu

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

AI总结 本文提出DiG框架,通过学习相似图像对的差异识别提升多模态大语言模型的细粒度感知能力,实验表明其在多个视觉感知基准上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04777 2026-01-09 cs.CV cs.AI 90%

GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models

GeM-VG:迈向通用多图像视觉 grounding 的多模态大语言模型

Shurong Zheng, Yousong Zhu, Hongyin Zhao, Fan Yang, Yufei Zhan, Ming Tang, Jinqiao Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 GeM-VG通过引入MG-Data-240K数据集和混合强化微调策略,提升多图像视觉 grounding 和通用多图像理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06118 2025-12-23 cs.CV cs.AI 90%

ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling

ViGoR:通过细粒度奖励建模提升大视觉语言模型的视觉 grounding

Siming Yan, Min Bai, Weifeng Chen, Xiong Zhou, Qixing Huang, Li Erran Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) AWS AI(AWS人工智能)

专题命中 视觉定位与Grounding :vision language model(title,abstract);grounding(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI

AI总结 ViGoR通过细粒度奖励建模提升大视觉语言模型的视觉 grounding 能力,采用更经济的人类评估和自动化方法,有效提高视觉推理准确性。

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06146 2025-11-11 cs.CL cs.AI cs.CV 90%

Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models

Akshar Tumu, Varad Shinde, Parisa Kordjamshidi

机构 * UC San Diego(加州大学圣地亚哥分校) IIT Kanpur(印度理工学院坎pur) Michigan State University(密歇根州立大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted at IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23903 2025-09-10 cs.CV cs.AI 90%

Grounding DINO-US-SAM: Text-Prompted Multi-Organ Segmentation in Ultrasound with LoRA-Tuned Vision-Language Models

Hamza Rasaee, Taha Koleilat, Hassan Rivaz

机构 * Department of Electrical and Computer Engineering, Concordia University(电气与计算机工程系,康科迪亚大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments 11 pages, 3 figures, 7 tables

Journal ref IEEE Transactions on Ultrasonics, Ferroelectrics, and Frequency Control, Sept. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20188 2025-08-29 cs.CV cs.LG 90%

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

Max Torop, Masih Eskandar, Nicholas Kurtansky, Jinyang Liu, Jochen Weber, Octavia Camps, Veronica Rotemberg, Jennifer Dy, Kivanc Kose

机构 * Northeastern University(东北大学) Memorial Sloan Kettering Cancer Center(纪念斯隆凯特琳癌症中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04702 2025-07-08 cs.CV cs.AI 90%

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning

Feng Yue, Zhaoxing Zhang, Junming Jiao, Zhengyu Liang, Shiwen Cao, Feifei Zhang, Rong Shen

专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08456 2026-04-10 cs.CV cs.CL 90%

Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models

熵-梯度基础:在视觉语言模型中无需训练的证据检索

Marcel Gröpl, Jaewoo Jung, Seungryong Kim, Marc Pollefeys, Sunghwan Hong

机构 * ETH Zurich(苏黎世联邦理工学院) ETH AI Center(ETH人工智能中心) KAIST AI(韩国科学技术院人工智能系)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(title,abstract);VLM(abstract);分类 cs.CV

AI总结 本文提出一种无需训练的视觉语言模型证据检索方法,通过熵梯度映射提升对微小视觉细节和多区域线索的处理能力,实验显示在多个基准上取得显著改进。

Comments Project Page : https://entropy-gradient-grounding.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13860 2024-10-18 cs.CV cs.RO 90%

VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Runsen Xu, Zhiwei Huang, Tai Wang, Yilun Chen, Jiangmiao Pang, Dahua Lin

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(title,abstract);vision-language model(abstract);分类 cs.CV

Comments CoRL 2024 Camera Ready. 25 pages. A novel zero-shot 3D visual grounding framework based solely on 2D images

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21744 2026-08-25 cs.HC 新提交 89%

CALM-BP: Observation-Matched Physiological Semantic Grounding for Non-Contact Blood Pressure Estimation

CALM-BP:用于非接触式血压估计的观测匹配生理语义 grounding

Haiyang Sun, Boyuan Gu, Yongjie Liu

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 提出 CALM-BP 方法,构建观测匹配生理语义 grounding,基于 FlowBP-Set 数据集验证语言可通过组织生理观测助力非接触式血压估计。

Comments 17 pages, 3 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17146 2026-08-19 cs.RO 新提交 89%

PDDL-ART: Autonomous Symbolic Abstraction From Demonstration For Long-Horizon Robotic Manipulation Using Vision-Language Models

PDDL-ART:基于演示的自主符号抽象方法,用于使用视觉语言模型的长时程机器人操纵

Disha Kamale, Dmitry Berenson

机构 * University of Michigan(密歇根大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title,abstract);grounding(abstract)

AI总结 PDDL-ART是一种基于VLM的框架,可自主从演示生成PDDL描述,经多阶段校正确保语义对齐,在发动机维护和家庭操纵任务上平均成功率达93.3%,优于基线规划器。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16004 2026-08-18 cs.IR 新提交 89%

LineageRAG: Harnessing GraphRAG by Constructing Evidence Lineages with Source Grounding

LineageRAG:通过构建带源 grounding 的证据谱系来利用 GraphRAG

Linyao Zheng, Xuhang Shi, Zhifang Mao, Sai Zhou, Shuaixian An, Xiuquan Hou, Jinze Li

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 LineageRAG 针对现有 GraphRAG 未明确证据发现与源 grounding 关联的问题,构建带源 grounding 的证据谱系,在三个多跳问答数据集上较基线实现多项指标提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04933 2026-08-06 cs.RO 新提交 89%

Mimir: A Neuro-Symbolic Memory System with Dynamic Grounding for Embodied Agents in Interactive Environments

Mimir:面向交互环境中具身智能体的、具备动态 grounding 的神经符号记忆系统

Haoming Xu, Zhenlin He, Hengyi Wang, Jiafeng Xu, Hao Dong

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 本文提出Mimir,一种分离世界与任务记忆并具备动态grounding的神经符号记忆系统,在EB-ALFRED、EB-Habitat等具身任务中显著提升了智能体的长程执行成功率。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28273 2026-06-29 cs.CL 新提交 89%

Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models

视觉默认,先验覆盖:视觉-语言模型中感知-知识冲突的因果机制

Niclas Lietzow, Danielle Bitterman, Carsten Eickhoff, William Rudman, Michal Golovanevsky

机构 * University of Tübingen(图宾根大学) Harvard University(哈佛大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title,abstract);grounding(abstract)

AI总结 通过激活修补和消融实验,发现VLM中视觉默认激活,而先验知识依赖少量因果注意力头(2.5-4.8%),形成不对称因果结构。

Comments 14 pages, 11 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29585 2026-05-29 cs.CL 89%

World Models in Words: Auditing Physical State-Transition Commitments in Vision-Language Models

语言中的世界模型:审计视觉语言模型中的物理状态转换承诺

Emmanuelle Bourigault

机构 * University of Oxford(牛津大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title,abstract);grounding(abstract)

AI总结 提出WMW框架,通过要求VLM输出结构化轨迹(初始状态、状态转换、结果状态和答案)并利用混合验证器检查模式有效性、状态基础、转换一致性和答案-轨迹兼容性,揭示仅评估最终答案所隐藏的物理推理失败。

Comments 8 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15281 2026-05-21 cs.SE 89%

Semantic Grounding of Digital Twin Metamodels Using RDF Graphs

基于 RDF 图的数字孪生元模型语义 grounding

Faima Abbasi, Jean-Sébastien Sottet, Cedric Pruski

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 本文提出了一种基于 RDF 图的数字孪生元模型语义 grounding 方法,通过设计多层数字孪生模型、将元模型提升为 RDF 图以及图基对齐方法 SSM-OM,实现了多层数字孪生的语义一致性与互操作性。

Comments Submitted to Conference, 15 pages excluding references, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03417 2026-05-15 cs.CL 89%

FactNet: A Billion-Scale Knowledge Graph for Multilingual Factual Grounding

FactNet:一个十亿级的知识图谱用于多语言事实 grounding

Yingli Shen, Wen Lai, Jie Zhou, Xueren Zhang, Yudong Wang, Kangyang Luo, Shuo Wang, Ge Gao, Alexander Fraser, Maosong Sun

机构 * Tsinghua University(清华大学) Technical University of Munich(慕尼黑技术大学) ModelBest Inc.(ModelBest公司) Minzu University of China(民族大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 FactNet通过结合17亿条维基数据语句和301亿条证据指针,构建了一个十亿级多语言知识图谱,提供事实 grounding 的评估框架,并验证了跨语言结构的知识迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02791 2026-04-28 cs.CL 89%

Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension

使对话 grounding 数据丰富:一种三层数据合成框架用于通用指称表达理解

Juexi Shao, Siyou Li, Yujian Gan, Chris Madge, Vanja Karan, Massimo Poesio

机构 * Queen Mary University of London(伦敦玛丽女王大学) Queen's University Belfast(贝尔法斯特女王大学) University of Vienna(维也纳大学) Utrecht University(乌得勒支大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 本文提出三层数据合成框架,通过平衡真实性和可控性,生成可扩展的对话条件 grounding 监督,提升指称表达理解的性能。

Journal ref ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2026, pp. 18142-18146

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12781 2026-08-19 cs.CV 版本更新 89%

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

超越正确性:混合思维多模态大语言模型(MLLM)的响应行为基准测试与对齐

Xinming Wang, Weinong Wang, Hongming Yang, Yansong Lin, Zheng Ruan, Shangpin Peng, Qiming Peng, Nan Qiao, Fengyuan Lu, Guoqing Ma, Marito Li, Songyang Zhang, Saiyong Yang, Han Hu, Yonglong Tian, Xu-Yao Zhang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Large Language Model Department, Tencent(腾讯大语言模型部) University of Electronic Science and Technology of China(电子科技大学) Hong Kong University of Science and Technology(香港科技大学) Zhongguancun Academy(中关村学院)

专题命中 视觉定位与Grounding :MLLM(title_cn,summary_cn);grounding(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 该研究针对混合思维 MLLM 的思维与非思维模式响应错位问题,构建 PatternEval 基准并开发 PatternRL 方法,可减轻跨模式错位且任务性能损失极小。

Comments 8 tables and 6figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03631 2026-08-05 cs.CV 新提交 89%

SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification

SEER:用于受控空间关系分类的自 grounding 证据接口

Feixiang Liu, Likun Wang, Qiang Qiu, Hui Xu, Huawei Shen, Xueqi Cheng

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);grounding(title_cn,abstract);分类 cs.CV

AI总结 本文提出针对冻结VLM的无训练推理时证据接口SEER,通过构建查询特定证据缓解空间关系分类错误,在多数据集上取得显著性能提升,证明该干预措施的有效性。

Comments 23 pages total, 2 figures. Code: https://github.com/SouthWinter/SEER

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18373 2026-04-14 cs.CV 89%

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models

MASS:面向视觉-语言模型中物理推理与理解的运动感知空间-时间 grounding

Xiyang Wu, Zongxia Li, Jihui Jin, Guangyao Shi, Gouthaman KV, Vishnu Raj, Nilotpal Sinha, Jingxi Chen, Fan Du, Dinesh Manocha

机构 * University of Maryland(马里兰大学) Dolby Laboratories(杜比实验室) University of Southern California(南加州大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(title);vision language model(abstract);VLM(abstract)

AI总结 MASS通过引入空间-时间信号提升视觉-语言模型对物理现象的理解能力,提出MASS-Bench基准测试集并验证模型在物理推理中的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23067 2026-03-26 cs.CV 89%

MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding

MLLM-HWSI: 一种用于分层全滑动图像理解的多模态大语言模型

Basit Alawode, Arif Mahmood, Muaz Khalifa Al-Radi, Shahad Albastaki, Asim Khan, Muhammad Bilal, Moshira Ali Abdalla, Mohammed Bennamoun, Sajid Javed

机构 * Department of Computer Science, Khalifa University of Science and Technology(卡利法科技大学计算机科学系) Information Technology University(信息技术大学) KAU(卡乌大学) University of the Western Australia(西澳大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(title,abstract);grounding(abstract);分类 cs.CV

AI总结 本文提出MLLM-HWSI,一种分层全滑动图像级多模态大语言模型,通过四级尺度对齐视觉特征与病理语言,提升解释性证据接地推理能力,在六个CPath任务上取得新SOTA结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06965 2026-03-13 cs.CV 89%

MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images

MedMO:为医学图像构建和理解多模态大语言模型

Ankan Deria, Komal Kumar, Adinath Madhavrao Dukre, Eran Segal, Salman Khan, Imran Razzak

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 MedMO是一种基于通用MLLM架构构建的医学多模态基础模型,通过多阶段训练提升跨模态和任务的性能,超越现有开源基线,在医学图像识别和报告生成中取得显著提升。

Comments 21 pages, 6 figures and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏