arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-30 至 2025-12-30 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 11 篇

2311.15759 2025-12-30 cs.CL cs.AI cs.CV 85%

Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs

视觉增强大语言模型:赋能大语言模型中的多模态知识存储与共享

Yunxin Li, Zhenyu Liu, Baotian Hu, Wei Wang, Yuxin Ding, Xiaochun Cao, Min Zhang

机构 * Research Institute of Computing and Intelligence(计算与智能研究 institute) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出MKS2方法,通过多模态知识存储与共享增强大语言模型的推理能力,提升其在物理和常识知识场景下的表现。

Comments 21 pages, 7 figures; Accepted by IEEE TIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23545 2025-12-30 cs.CV cs.AI 81%

PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis

PathFound: 一种促进证据寻求病理诊断的代理多模态模型

Shengyi Hua, Jianfeng Wu, Tianle Shen, Kangzhe Hu, Zhongzhen Huang, Shujuan Ni, Zhihong Zhang, Yuan Li, Zhe Wang, Xiaofan Zhang

机构 * Qing Yuan Research Institute, Shanghai Jiao Tong University(上海交通大学庆元研究院) Shanghai Innovation Institute(上海创新研究院) Department of Pathology, The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine, University of Science and Technology of China(中国科学技术大学生命科学与医学学院病理科) Intelligent Pathology Institute, Division of Life Sciences and Medicine(生命科学与医学学院智能病理研究所) Department of Pathology, Fudan University Shanghai Cancer Center(复旦大学上海癌症中心病理科) Department of Oncology, Shanghai Medical College, Fudan University(复旦大学上海医学院肿瘤科) Institute of Pathology, Fudan University(复旦大学病理研究所) Department of Pathology, The First Affiliated Hospital with Nanjing Medical University(南京医科大学第一附属医院病理科)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 PathFound是一种通过主动信息获取和诊断完善来提升病理诊断准确性的代理多模态模型,其在多种临床场景中表现出卓越的诊断性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23573 2025-12-30 cs.CV 79%

ProGuard: Towards Proactive Multimodal Safeguard

ProGuard:迈向主动多模态安全防护

Shaohan Yu, Lijun Li, Chenyang Si, Lu Sheng, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) PRLab Nanjing University(南京大学PRLab) Beihang University(北航)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 ProGuard通过强化学习和同义词库奖励机制,实现主动多模态安全防护,显著提升分布外风险检测与描述能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23557 2025-12-30 cs.CR cs.AI 79%

Toward Trustworthy Agentic AI: A Multimodal Framework for Preventing Prompt Injection Attacks

迈向可信的代理AI:一种多模态框架用于防止提示注入攻击

Toqeer Ali Syed, Mishal Ateeq Almutairi, Mahmoud Abdel Moaty

机构 * Faculty of Computer and Information System(计算机与信息系统系) Islamic University of Madinah(麦地那伊斯兰大学) Arab Open University-Bahrain(巴林阿拉伯开放大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出一种多模态框架,通过溯源感知机制防止代理AI中的提示注入攻击,提升系统安全性和稳定性。

Comments It is accepted in a conference paper, ICCA 2025 in Bahrain on 21 to 23 December

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23453 2025-12-30 cs.CV cs.AI 62%

CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models

CoFi-Dec:通过粗到细的生成反馈在大视觉-语言模型中实现抗幻觉解码

Zongsheng Cao, Yangfan He, Anran Liu, Jun Xie, Feng Chen, Zepeng Wang

机构 * Researcher(研究者)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 CoFi-Dec通过粗到细的生成反馈机制,在大视觉-语言模型中减少幻觉,提升解码的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23244 2025-12-30 cs.CV cs.AI 62%

ViLaCD-R1: A Vision-Language Framework for Semantic Change Detection in Remote Sensing

ViLaCD-R1: 一种用于遥感语义变化检测的视觉-语言框架

Xingwei Ma, Shiyang Feng, Bo Zhang, Bin Wang

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 ViLaCD-R1通过多图像推理器和掩码引导解码器,提升遥感变化检测的语义识别与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23028 2025-12-30 cs.CV cs.AI cs.SE 62%

An Architecture-Led Hybrid Report on Body Language Detection Project

以架构为导向的混合报告:身体语言检测项目

Thomson Tong, Diba Darooneh

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本报告分析了两种视觉-语言模型的架构特性,并探讨了其在视频到制品管道中的应用及系统限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22238 2025-12-30 cs.LG cs.AI cs.CV 62%

Masking Teacher and Reinforcing Student for Distilling Vision-Language Models

掩码教师与强化学生用于蒸馏视觉-语言模型

Byung-Kwan Lee, Yu-Chiang Frank Wang, Ryo Hachiuma

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Masters通过掩码教师和强化学生的方法,解决视觉-语言模型蒸馏中的大小差距问题,提升学生模型的表示学习能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06474 2025-12-30 cs.CV cs.AI cs.LG 62%

Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding

通过不确定性引导的丢弃解码增强视觉-语言模型的可靠性

Yixiong Fang, Ziran Yang, Zhaorun Chen, Zhuokai Zhao, Jiawei Zhou

机构 * Carnegie Mellon University(卡内基梅隆大学) Princeton University(普林斯顿大学) University of Chicago(芝加哥大学) Stony Brook University(石溪大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 通过不确定性引导的丢弃解码方法,有效减少视觉-语言模型的幻觉问题,提升输出的可靠性和质量。

Comments Accepted to 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23109 2025-12-30 cs.LG cs.AI stat.ML 57%

How Much Data Is Enough? Uniform Convergence Bounds for Generative & Vision-Language Models under Low-Dimensional Structure

需要多少数据?在低维结构下生成式与视觉-语言模型的统一收敛界限

Paul M. Thompson

机构 * Stevens Institute for Neuroimaging and Informatics, University of Southern California(神经影像与信息学研究所,南加州大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.AI

AI总结 研究在低维结构下生成式和视觉-语言模型的统一收敛界限,探讨数据量与模型校准之间的关系。

Comments 13 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16531 2025-12-30 cs.AI 57%

Scaling Laws for Energy Efficiency of Local LLMs

本地大语言模型和视觉-语言模型的扩展定律

Ander Alvarez, Alessandro Genuardi, Nilotpal Sinha, Antonio Tiene, Mikail Okyay, Bakbergen Ryskulov, David Montero, Samuel Mugel, Román Orús

机构 * Multiverse Computing Donostia International Physics Center Ikerbasque Foundation for Science Multiverse Computing, Centre for Social Innovation

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

AI总结 研究揭示了本地大语言模型和视觉-语言模型在中央处理单元上的计算扩展定律,并展示了量子启发压缩在降低能耗和资源消耗方面的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏