arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-16 至 2026-02-16 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 8 篇

2602.12936 2026-02-16 cs.CV 85%

Unleashing MLLMs on the Edge: A Unified Framework for Cross-Modal ReID via Adaptive SVD Distillation

在边缘侧释放MLLMs:通过自适应SVD蒸馏实现跨模态重识别的统一框架

Hongbo Jiang, Jie Li, Xinqi Cai, Tianyu Xie, Yunhang Shen, Pingyang Dai, Liujuan Cao

机构 * Tencent Youtu Lab, Shanghai, China(腾讯优图实验室,上海,中国) Xiamen University, Media Analytics(厦门大学,媒体分析) Computing Lab, Department of Artificial Intelligence, School of Informatics, Xiamen, China(计算实验室,人工智能系,信息学院,厦门,中国)

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出MLLMEmbed-ReID框架,通过自适应SVD蒸馏实现跨模态重识别的统一部署,结合云-边缘架构提升性能与效率。

Comments Equal contribution by Jie Li

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07909 2026-02-16 cs.LG cs.AI cs.CV 84%

Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning

解释并缓解对比多模态学习中的模态差距

Can Yaras, Siyi Chen, Peng Wang, Qing Qu

机构 * Department of Electrical Engineering \& Computer Science, University of Michigan

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本文通过分析对比多模态学习中的模态差距成因,提出通过温度调度和模态交换等策略缓解模态差距,从而提升多模态任务性能。

Comments The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10271 2026-02-16 cs.IR cs.CL 83%

MLDocRAG: Multimodal Long-Context Document Retrieval Augmented Generation

MLDocRAG: 多模态长上下文文档检索增强生成

Yongyue Zhang, Yaxiong Wu

机构 * Independent Researcher(独立研究者)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 MLDocRAG通过构建多模态片段-查询图,实现多模态长上下文文档的检索增强生成,提升问答的准确性和连贯性。

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01390 2026-02-16 cs.CV cs.AI cs.LG 81%

Multimodal Doctor-in-the-Loop: A Clinically-Guided Explainable Framework for Predicting Pathological Response in Non-Small Cell Lung Cancer

多模态医生在环:一种临床指导的可解释框架,用于预测非小细胞肺癌的病理反应

Alice Natalina Caragliano, Claudia Tacconi, Carlo Greco, Lorenzo Nibid, Edy Ippolito, Michele Fiore, Giuseppe Perrone, Sara Ramella, Paolo Soda, Valerio Guarrasi

机构 * Research Unit of Computer Systems and Bioinformatics, Department of Engineering(计算机系统与生物信息学研究单位,工程系) Research Unit of Radiation Oncology, Department of Medicine and Surgery(放射肿瘤学研究单位,医学与外科系) Research Unit of Anatomical Pathology, Department of Medicine and Surgery(病理学研究单位,医学与外科系)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究提出多模态医生在环框架,结合影像与临床数据,通过嵌入医生知识提升肺癌病理反应预测的准确性和可解释性。

Comments arXiv admin note: substantial text overlap with arXiv:2502.17503

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12774 2026-02-16 cs.CV 79%

Bootstrapping MLLM for Weakly-Supervised Class-Agnostic Object Counting

基于MLLM的弱监督类无关物体计数

Xiaowen Zhang, Zijie Yue, Yong Luo, Cairong Zhao, Qijun Chen, Miaojing Shi

机构 * College of Electronic and Information Engineering, Tongji University(电子信息工程学院,同济大学) College of Computer Science, Wuhan University(计算机科学学院,武汉大学) College of Computer Science, Tongji University(计算机科学学院,同济大学) State Key Laboratory of Autonomous Intelligent Unmanned Systems(自主智能无人系统国家重点实验室)

专题命中 跨模态检索 :MLLM(title,abstract);分类 cs.CV

AI总结 本文提出 WS-COC,基于 MLLM 的弱监督框架,通过三种策略提升类无关物体计数性能,有效降低标注成本。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05096 2026-02-16 cs.CV cs.LG 79%

Visual concept ranking uncovers medical shortcuts used by large multimodal models

视觉概念排名揭示大型多模态模型使用的医学捷径

Joseph D. Janizek, Sonnet Xu, Junayd Lateef, Roxana Daneshjou

机构 * Stanford University(斯坦福大学) University of California Berkeley(加州大学伯克利分校)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出视觉概念排名方法,用于揭示大型多模态模型在医疗任务中表现的潜在视觉特征依赖性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13179 2026-02-16 cs.IR 78%

Fix Before Search: Benchmarking Agentic Query Visual Pre-processing in Multimodal Retrieval-augmented Generation

在搜索前修复:多模态检索增强生成中代理查询视觉预处理的基准测试

Jiankun Zhang, Shenglai Zeng, Kai Guo, Xinnan Dai, Hui Liu, Jiliang Tang, Yi Chang

专题命中 跨模态检索 :multimodal(title,abstract)

AI总结 本文提出V-QPP-Bench,通过代理决策任务评估视觉查询预处理的改进,揭示视觉不完美对检索性能的严重影响及训练方法的优化潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12883 2026-02-16 eess.IV cs.CV 74%

Dual-Phase Cross-Modal Contrastive Learning for CMR-Guided ECG Representations for Cardiovascular Disease Assessment

双相跨模态对比学习用于CMR引导的ECG表示以评估心血管疾病

Laura Alvarez-Florez, Angel Bujalance-Gomez, Femke Raijmakers, Samuel Ruiperez-Campillo, Maarten Z. H. Kolk, Jesse Wiers, Julia Vogt, Erik J. Bekkers, Ivana Išgum, Fleur V. Y. Tjong

机构 * Amsterdam University Medical Center(阿姆斯特丹大学医学中心) University of Amsterdam(阿姆斯特丹大学) ETH Zurich(苏黎世联邦理工学院) Department of Radiology and Nuclear Medicine, Amsterdam University Medical Center(放射医学与核医学系,阿姆斯特丹大学医学中心) Department of Radiology, Mayo Clinic, Rochester(放射医学系,梅奥诊所,罗切斯特)

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

AI总结 本文提出双相跨模态对比学习方法,通过结合ECG和CMR数据提升心血管疾病评估的ECG表示能力。

Comments Paper accepted at SPIE Medical Imaging 2026 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏