arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-17 至 2025-12-17 共收录 57 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2412.07148 2025-12-17 cs.CV cs.AI cs.CL cs.LG 82%

MM-PoE: Multiple Choice Reasoning via. Process of Elimination using Multi-Modal Models

MM-PoE:通过多模态模型的排除过程进行多选推理

Sayak Chakrabarty, Souradip Pal

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 MM-PoE通过多模态模型的排除过程提升视觉语言模型在多选推理任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13747 2025-12-17 cs.CV cs.AI 82%

Why Text Prevails: Vision May Undermine Multimodal Medical Decision Making

为何文本占上风:视觉可能损害多模态医疗决策制定

Siyuan Dai, Lunxiao Li, Kun Zhao, Eardi Lila, Paul K. Crane, Heng Huang, Dongkuan Xu, Haoteng Tang, Liang Zhan

机构 * University of Texas Rio Grande Valley(德克萨斯大学里奥格兰德谷大学) University of Pittsburgh(匹兹堡大学) NC State University(北卡罗来纳州立大学) University of Washington(华盛顿大学) University of Maryland(马里兰大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究发现文本推理在医疗多模态决策中优于多模态输入,提出三种策略以提升多模态医疗决策能力。

Comments Accepted by ICDM 2025 the Workshop on Synergy of AI and Multimodal Biomedical Data Mining

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14050 2025-12-17 cs.CV 70%

SELECT: Detecting Label Errors in Real-world Scene Text Data

SELECT:在现实世界场景文本数据中检测标签错误

Wenjun Liu, Qian Wu, Yifeng Hu, Yuke Li

机构 * Yidun AI Lab, NetEase(网易异度人工智能实验室)

专题命中 图文多模态 :multi-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 SELECT通过多模态训练和SSLC方法有效检测现实世界场景文本数据中的标签错误,提升文本识别准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20047 2025-12-17 cs.CV eess.IV 70%

Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image Analysis

Med3DVLM: 一种高效的视觉-语言模型用于3D医学图像分析

Yu Xin, Gorkem Can Ates, Kuang Gong, Wei Shao

机构 * University of Florida(佛罗里达大学)

专题命中 图文多模态 :multi-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 Med3DVLM通过三个创新提出,实现了高效的3D医学图像分析,显著提升了图像-文本检索、报告生成和视觉问答的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14480 2025-12-17 cs.CV 57%

SuperCLIP: CLIP with Simple Classification Supervision

SuperCLIP:带有简单分类监督的CLIP

Weiheng Zhao, Zilong Huang, Jiashi Feng, Xinggang Wang

机构 * School of EIC, Huazhong University of Science and Technology(华中科技大学电子与信息学院) ByteDance(字节跳动)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

AI总结 SuperCLIP通过添加轻量级分类监督提升CLIP的细粒度视觉-文本对齐能力,有效提升零样本分类和检索性能。

Comments Accepted by NeurIPS 2025. Code: https://github.com/hustvl/SuperCLIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14137 2025-12-17 cs.CV 57%

Erasing CLIP Memories: Non-Destructive, Data-Free Zero-Shot class Unlearning in CLIP Models

擦除CLIP记忆:非破坏性、无数据零样本类去学习在CLIP模型中

Ashish Mishra, Tarun Kumar, Gyanaranjan Nayak, Arpit Shah, Suparna Bhattacharya, Martin Foltin

机构 * Hewlett Packard Labs(惠普实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种基于空空间投影的非破坏性零样本类去学习方法,有效降低CLIP模型中目标类别的性能,同时保留整体多模态知识。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14113 2025-12-17 cs.CV 57%

Selective, Controlled and Domain-Agnostic Unlearning in Pretrained CLIP: A Training- and Data-Free Approach

选择性、可控性和领域无关的预训练CLIP模型去学习:一种无需训练和数据的方法

Ashish Mishra, Gyanaranjan Nayak, Tarun Kumar, Arpit Shah, Suparna Bhattacharya, Martin Foltin

机构 * Hewlett Packard Labs(惠普实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种无需训练和数据的CLIP去学习方法,能够实现选择性、可控性和领域无关的模型遗忘,提升模型在不同任务上的适应性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02503 2025-12-17 cs.CV 57%

Adapting General-Purpose Foundation Models for X-ray Ptychography in Low-Data Regimes

为低数据情形下的X射线衍射成像适应通用基础模型

Robinson Umeike, Neil Getty, Yin Xiangyu, Yi Jiang

机构 * The University of Alabama(阿拉巴马大学) Argonne National Laboratory(阿贡国家实验室)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出PtychoBench基准,通过比较SFT和ICL策略,在低数据环境下优化X射线衍射成像任务的模型适应性,发现任务模态决定最佳专门化路径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14661 2025-12-17 cs.AR 50%

Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models

聚焦:一种高效的视觉-语言模型流式集中架构

Chiyue Wei, Cong Guo, Junyao Zhang, Haoxuan Shan, Yifan Xu, Ziyue Zhang, Yudong Liu, Qinsi Wang, Changchun Zhou, Hai "Helen" Li, Yiran Chen

专题命中 图文多模态 :cross-modal(abstract)

AI总结 Focus提出了一种高效的视觉-语言模型流式集中架构,通过分层压缩和细粒度冗余消除,实现2.4倍速度提升和3.3倍能效提升。

Comments HPCA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 3 篇

2512.14083 2025-12-17 eess.AS cs.CL cs.LG 84%

Scalable Frameworks for Real-World Audio-Visual Speech Recognition

可扩展的音频视觉语音识别系统框架

Sungnyun Kim

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CL、eess.AS

AI总结 本文提出了一种可扩展的音频视觉语音识别系统框架,通过分层方法提升模型在现实环境中的鲁棒性和可扩展性。

Comments PhD Dissertation

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14115 2025-12-17 cs.SD cs.LG 82%

Joint Multimodal Contrastive Learning for Robust Spoken Term Detection and Keyword Spotting

多模态对比学习联合框架用于鲁棒的语音术语检测和关键词 spotting

Ramesh Gundluru, Shubham Gupta, Sri Rama Murty K

机构 * Electrical Engineering(电子工程) Artificial Intelligence(人工智能)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文提出一种联合多模态对比学习框架,通过统一音频和跨模态监督提升语音术语检测和关键词 spotting 的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13981 2025-12-17 cs.RO 50%

Impact of Robot Facial-Audio Expressions on Human Robot Trust Dynamics and Trust Repair

机器人面部-语音表达对人类机器人信任动态及信任修复的影响

Hossein Naderi, Alireza Shojaei, Philip Agee, Kereshmeh Afsari, Abiola Akanmu

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 本研究探讨了机器人面部-语音表达对人类信任动态及修复的影响,发现成功提升信任,失败导致信任下降,道歉表达部分恢复信任,且年龄和先前态度影响信任变化的持续时间。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 4 篇

2503.09143 2025-12-17 cs.CV 83%

Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding

Exo2Ego:基于外部知识引导的多模态大语言模型用于第一人称视频理解

Haoyu Zhang, Qiaohui Chu, Meng Liu, Haoxiang Shi, Yaowei Wang, Liqiang Nie

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 Exo2Ego通过迁移学习提升内向视频理解能力,利用外向知识增强模型性能。

Comments This paper is accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14058 2025-12-17 cs.CV cs.AI 81%

Real-time prediction of workplane illuminance distribution for daylight-linked controls using non-intrusive multimodal deep learning

基于非侵入式多模态深度学习的实时工作平面照度分布预测用于日光联动控制

Zulin Zhuang, Yu Bian

机构 * School of Architecture(建筑学院) State Key Laboratory of Subtropical Building(亚热带建筑科学国家重点实验室) South China University of Technology(华南理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究提出了一种基于非侵入式多模态深度学习的实时工作平面照度分布预测方法,用于提高日光联动控制的能效

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10021 2025-12-17 cs.LG cs.AI 79%

Online Multi-modal Root Cause Identification in Microservice Systems

微服务系统中的在线多模态根本原因识别

Lecheng Zheng, Zhengzhang Chen, Haifeng Chen

机构 * University of Illinois Urbana-Champagin(伊利诺伊大学厄巴纳-香槟分校) NEC Labs America(NEC美洲实验室)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出OCEAN方法,通过多模态因果结构学习实现微服务系统中的在线根本原因识别,结合扩张卷积神经网络和图神经网络,提升因果图学习的准确性与效率。

Comments Accepted by BigData 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14017 2025-12-17 cs.CV cs.AI 62%

KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding

KFS-Bench: 长视频理解中关键帧采样的全面评估

Zongyao Li, Kengo Ishida, Satoshi Yamazaki, Xiaotong Ji, Jianquan Liu

机构 * Visual Intelligence Research Laboratories, NEC Corporation(NEC公司视觉智能研究实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 KFS-Bench通过多场景标注评估关键帧采样策略,提出新度量标准和方法提升问答性能。

Comments WACV2026

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 5 篇

2511.15464 2025-12-17 cs.CV cs.LG 83%

SIGMMA: Hierarchical Graph-Based Multi-Scale Multi-modal Contrastive Alignment of Histopathology Image and Spatial Transcriptome

SIGMMA:基于层次图的多尺度多模态对比对齐:组织病理图像与空间转录组

Dabin Jeong, Amirhossein Vahidi, Ciro Ramírez-Suástegui, Marie Moullet, Kevin Ly, Mohammad Vali Sanian, Sebastian Birk, Yinshui Chang, Adam Boxall, Daniyal Jafree, Lloyd Steele, Vijaya Baskar MS, Muzlifah Haniffa, Mohammad Lotfollahi

机构 * Wellcome Sanger Institute(沃森桑格研究所) Cambridge Centre for AI in Medicine(剑桥人工智能医学中心) Institute of AI for Health(人工智能与健康研究所) Cambridge Stem Cell Institute(剑桥干细胞研究所)

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 SIGMMA通过多尺度多模态对比对齐,提升组织病理图像与空间转录组的跨模态对应表示,提高基因表达预测和跨模态检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14230 2025-12-17 cs.LG stat.ML 78%

Understanding the Gain from Data Filtering in Multimodal Contrastive Learning

理解数据过滤在多模态对比学习中的收益

Divyansh Pareek, Sewoong Oh, Simon S. Du

机构 * Paul G. Allen School of Computer Science and Engineering(保罗·G·艾伦计算机科学与工程学院)

专题命中 跨模态检索 :multimodal(title,abstract)

AI总结 本文研究了数据过滤在多模态对比学习中的收益,通过理论分析证明了过滤能有效降低误差,提升模型性能。

Comments 40 pages, 8 figures, 1 table. This work is accepted to the Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13164 2025-12-17 cs.CV cs.AI 62%

A Semantically Enhanced Generative Foundation Model Improves Pathological Image Synthesis

语义增强的生成基础模型提升病理图像合成

Xianchao Guan, Zhiyuan Fan, Yifeng Wang, Fuqiang Chen, Yanjiang Zhou, Zengyang Che, Hongxue Meng, Xin Li, Yaowei Wang, Hongpeng Wang, Min Zhang, Heng Tao Shen, Zheng Zhang, Yongbing Zhang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 CRAFTS通过语义增强的生成基础模型提升病理图像合成质量,生成多样化病理图像并增强多种临床任务性能。

Comments 68 pages, 9 figures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16313 2025-12-17 cs.LG cs.AI cs.CL 62%

Retrieval Enhanced Feedback via In-context Neural Error-book

通过上下文神经错误本增强的检索反馈

Jongyeop Hyun, Bumsoo Kim

机构 * School of CSE Chung-Ang University(计算机科学与工程学院 Chung-Ang 大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 REFINE通过结构化反馈框架,系统分析并缓解多模态推理中的错误,提升推理效率和可扩展性。

Comments Accepted at EMNLP 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01115 2025-12-17 cs.AI cs.MA econ.TH physics.soc-ph 57%

Exploring Network-Knowledge Graph Duality: A Case Study in Agentic Supply Chain Risk Analysis

探索网络-知识图谱二元性:代理供应链风险分析的案例研究

Evan Heus, Rick Bookstaber, Dhruv Sharma

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出基于LLM的代理框架,利用网络与知识图谱的二元性,通过图遍历和上下文壳技术实现供应链风险分析的实时可解释生成。

Comments Accepted to the 2nd Workshop on LLMs and Generative AI in Finance: International Conference on AI in Finance(ICAIF) 2025;7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 5 篇

2512.13752 2025-12-17 cs.CV cs.AI 81%

STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning

STAR:用于统一多模态学习的堆叠自回归方案

Jie Qin, Jiancheng Huang, Limeng Qiao, Lin Ma

机构 * Meituan Inc(美团公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 STAR通过堆叠自回归方案提升多模态生成性能,同时保持理解能力,实验验证其在统一多模态学习中的有效性。

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14225 2025-12-17 cs.CV 79%

OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving

OmniGen:面向自动驾驶的统一多模态传感器生成

Tao Tang, Enhui Ma, xia zhou, Letian Wang, Tianyi Yan, Xueyang Zhang, Kun Zhan, Peng Jia, XianPeng Lang, Jia-Wang Bian, Kaicheng Yu, Xiaodan Liang

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Westlake University(西湖大学) Li Auto Inc.(Li汽车公司) The University of Toronto(多伦多大学) University of Macau(澳门大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 OmniGen通过统一框架生成对齐的多模态传感器数据,结合共享BEV空间和UAE方法,实现高效且灵活的传感器生成。

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13729 2025-12-17 cs.LG cs.AI cs.CV 76%

Composite Classifier-Free Guidance for Multi-Modal Conditioning in Wind Dynamics Super-Resolution

多模态条件下的风动力超分辨率复合分类器引导方法

Jacob Schnell, Aditya Makkar, Gunadi Gani, Aniket Srinivasan Ashok, Darren Lo, Mike Optis, Alexander Wong, Yuhao Chen

机构 * University of Waterloo(滑铁卢大学) Veer Renewables

专题命中 多模态生成 :multi-modal(title);分类 cs.CV、cs.AI

AI总结 本文提出复合分类器引导方法,用于多模态条件下的风动力超分辨率重建,实现高保真度与低成本的风数据生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14217 2025-12-17 cs.CV cs.RO 57%

DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos

DRAW2ACT:将深度编码轨迹转化为机器人演示视频

Yang Bai, Liudi Yang, George Eskandar, Fengyi Shen, Mohammad Altillawi, Ziyuan Liu, Gitta Kutyniok

机构 * Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) University of Freiburg(弗赖堡大学) Technical University of Munich(慕尼黑技术大学) Huawei Heisenberg Research Center (Munich)(华为海森堡研究中心(慕尼黑))

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 DRAW2ACT通过深度感知的轨迹条件视频生成框架,生成可控且一致的机器人演示视频,提升机器人操作的成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11252 2025-12-17 cs.CV eess.IV 57%

MFGDiffusion: Mask-Guided Smoke Synthesis for Enhanced Forest Fire Detection

MFGDiffusion:基于掩码的烟雾合成以提升森林火灾检测

Guanghao Wu, Yunqing Shang, Chen Xu, Hai Song, Chong Wang, Qixing Zhang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 MFGDiffusion通过掩码引导的烟雾合成提升森林火灾检测性能

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 14 篇

2512.14594 2025-12-17 cs.CV 83%

LLM-driven Knowledge Enhancement for Multimodal Cancer Survival Prediction

基于大语言模型的知识增强的多模态癌症生存预测

Chenyu Zhao, Yingxue Xu, Fengtao Zhou, Yihui Wang, Hao Chen

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology(香港科技大学化学与生物工程系) Division of Life Science, The Hong Kong University of Science and Technology(香港科技大学生命科学系) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology(香港科技大学深圳-香港协同创新研究院) State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology(香港科技大学神经系统疾病国家重点实验室)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出KEMM模型,通过整合专家报告和预后背景知识,利用知识增强的跨模态注意力模块提升癌症生存预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14620 2025-12-17 cs.CL cs.AI cs.CV 82%

JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction

JMMMU-Pro:通过Vibe基准构建的基于图像的日本多学科多模态理解基准

Atsuyuki Miyai, Shota Onohara, Jeonghun Baek, Kiyoharu Aizawa

机构 * The University of Tokyo(东京大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 JMMMU-Pro通过构建基于图像的多模态理解基准,评估LMMs在日文处理能力,提出Vibe基准构建方法以提高基准质量。

Comments Project page: https://mmmu-japanese-benchmark.github.io/JMMMU_Pro/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14491 2025-12-17 cs.AI 79%

Sparse Multi-Modal Transformer with Masking for Alzheimer's Disease Classification

用于阿尔茨海默病分类的稀疏多模态Transformer与掩码

Cheng-Han Lu, Pei-Hsuan Tsai

机构 * Cheng-Han Lu(无) Pei-Hsuan Tsai(无)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出SMMT,一种稀疏多模态Transformer架构,通过稀疏注意力和模态掩码提升效率与鲁棒性,应用于阿尔茨海默病分类任务。

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01728 2025-12-17 cs.CV 79%

Multimodal classification of forest biodiversity potential from 2D orthophotos and 3D airborne laser scanning point clouds

基于2D正射影像和3D空中激光扫描点云的森林生物多样性潜力多模态分类

Simon B. Jensen, Stefan Oehmcke, Andreas Møgelmose, Meysam Madadi, Christian Igel, Sergio Escalera, Thomas B. Moeslund

机构 * Perception Laboratory, Aalborg University, Denmark Pioneer Centre for Artificial Intelligence, Denmark Department of Computer Science, Copenhagen University, Denmark Institute for Visual \& Analytic Computing, Rostock University, Germany University of Barcelona Computer Vision Center, Spain

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究利用2D正射影像和3D ALS点云数据,通过多模态深度学习融合方法,实现对森林生物多样性潜力的高效评估,实验结果达到82%的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏