arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-23 至 2026-01-23 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 6 篇

2405.10739 2026-01-23 cs.CV cs.AI 84%

Efficient Multimodal Large Language Models: A Survey

高效多模态大语言模型:综述

Yizhang Jin, Jian Li, Yexin Liu, Tianjun Gu, Kai Wu, Zhengkai Jiang, Muyang He, Bo Zhao, Xin Tan, Zhenye Gan, Yabiao Wang, Chengjie Wang, Lizhuang Ma

机构 * Youtu Lab, Tencent(腾讯优图实验室) SJTU(上海交通大学) BAAI(北京人工智能研究院) ECNU(华东师范大学)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 本文综述了高效多模态大语言模型的发展现状,探讨了其高效结构、策略及应用,并展望了未来研究方向。

Comments Accepted by Visual Intelligence

Journal ref Visual Intelligence, Volume 3, article number 27, (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11304 2026-01-23 cs.AI cs.CL cs.CV 82%

Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring

利用实例分割辅助的多模态大语言模型进行智能交通监控

Murat Arda Onsu, Poonam Lohan, Burak Kantarci, Aisha Syed, Matthew Andrews, Sean Kennedy

机构 * University of Ottawa(渥太华大学) Nokia Bell Labs(诺基亚贝尔实验室)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文利用多模态大语言模型和实例分割技术,实现高准确率的交通监控系统,提升交通管理效率和安全性。

Comments 6 pages, 7 figures, submitted to 30th IEEE International Symposium on Computers and Communications (ISCC) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15947 2026-01-23 physics.med-ph physics.app-ph physics.bio-ph physics.ins-det physics.optics 78%

Multimodal Imaging System Combining Hyperspectral and Laser Speckle Imaging for In Vivo Hemodynamic and Metabolic Monitoring

多模态成像系统:结合超光谱成像与激光散斑成像用于活体血流和代谢监测

Junda Wang, Luca Giannoni, Ayse Gertrude Yenicelik, Eleni Giama, Frederic Lange, Kenneth J. Smith, Ilias Tachtsidis

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本研究提出了一种结合超光谱和激光散斑成像的多模态成像系统,用于实时监测活体组织的血流和代谢状态。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15347 2026-01-23 cs.AI cs.CL cs.LG 62%

Logic Programming on Knowledge Graph Networks And its Application in Medical Domain

知识图谱网络上的逻辑编程及其在医疗领域的应用

Chuanqing Wang, Zhenmin Zhao, Shanshan Du, Chaoqun Fei, Songmao Zhang, Ruqian Lu

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出知识图谱网络的系统理论与技术,探讨其在医疗领域的应用,通过多条件下的实验验证创新方法。

Comments 33 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14625 2026-01-23 cs.CR cs.AI 57%

VTarbel: Targeted Label Attack with Minimal Knowledge on Detector-enhanced Vertical Federated Learning

VTarbel:基于最小知识的目标标签攻击与增强垂直联邦学习

Juntao Tan, Anran Li, Quanchao Liu, Peng Ran, Lan Zhang

机构 * University of Science and Technology of China(科学技术大学) Department of Security Technology Research, China Mobile Research Institute(安全技术研究所)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 VTarbel是一种针对增强垂直联邦学习的最小知识目标标签攻击框架,通过两阶段方法规避检测并有效诱导误分类。

Comments Accepted by ACM Transactions on Sensor Networks (TOSN)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15716 2026-01-23 cs.CE cs.CR 50%

zkFinGPT: Zero-Knowledge Proofs for Financial Generative Pre-trained Transformers

zkFinGPT: 用于金融生成预训练变换器的零知识证明

Xiao-Yang Liu, Ningjie Li, Keyi Wang, Xiaoli Zhi, Weiqin Tong

专题命中 其他多模态 :multimodal(abstract)

AI总结 zkFinGPT通过零知识证明技术,在保护数据隐私的同时验证金融生成预训练变换器的合法性与可信度,但其较高的计算开销限制了实际应用。

详情

展开后加载摘要…

URL PDF HTML 收藏