arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-01 至 2025-12-01 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 6 篇

2511.21827 2025-12-01 cs.AI cs.CV 84%

Evaluating Strategies for Synthesizing Clinical Notes for Medical Multimodal AI

评估合成临床笔记的策略以用于医疗多模态AI

Niccolo Marini, Zhaohui Liang, Sivaramakrishnan Rajaraman, Zhiyun Xue, Sameer Antani

机构 * Division of Intramural Research, National Library of Medicine, National Institutes of Health(国家医学图书馆内部研究部,国家卫生研究院)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文研究了合成临床笔记的生成策略,通过优化提示设计和医学元数据,提升医疗多模态AI在分类和跨模态检索任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21707 2025-12-01 cs.NI cs.AI 83%

Sensing and Understanding the World over Air: A Large Multimodal Model for Mobile Networks

通过空气感知和理解世界:为移动网络设计的大型多模态模型

Zhuoran Duan, Yuhao Wei, Guoshun Nan, Zijun Wang, Yan Yan, Lihua Xiong, Yuhan Ran, Ji Zhang, Jian Li, Qimei Cui, Xiaofeng Tao, Tony Q. S. Quek

机构 * National Engineering Research Center for Mobile Network Technologies, Beijing University of Posts and Telecommunications (BUPT), Beijing(中国移动网络技术国家工程研究中心,北京邮电大学) Beiyou Shenzhen Institute(北邮深圳研究所) School of Cyber Security, University of Chinese Academy of Sciences (UCAS)(中国科学院大学网络安全学院) China Telecom Co., Ltd.(中国电信股份有限公司) Singapore University of Technology and Design (SUTD)(新加坡科技设计大学)

专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

AI总结 本文提出了一种无线原生多模态模型,利用无线信号进行对比学习,验证了其在无线网络中的应用潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22997 2025-12-01 cs.CV cs.RO 79%

MrGS: Multi-modal Radiance Fields with 3D Gaussian Splatting for RGB-Thermal Novel View Synthesis

MrGS: 多模态辐射场与3D高斯点云融合用于RGB-热成像新视角合成

Minseong Kweon, Janghyun Kim, Ukcheol Shin, Jinsun Park

机构 * Minnesota Robotics Institute (MnRI), University of Minnesota, Twin Cities(明尼苏达大学罗学院(MnRI)、明尼苏达大学双城分校) Department of Information Convergence Engineering (Artificial Intelligence Major), Pusan National University(信息融合工程系(人工智能专业),釜山国立大学) Department of Energy Engineering, Korea Institute of Energy Technology (KENTECH)(能源工程系,韩国能源技术研究所(KENTECH)) School of Computer Science and Engineering, Pusan National University(计算机科学与工程学院,釜山国立大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 MrGS通过多模态辐射场与3D高斯点云融合,实现RGB和热成像新视角合成,利用物理原理建模热传导和辐射现象,提升重建精度与效率。

Comments Accepted at Thermal Infrared in Robotics (TIRO) Workshop, ICRA 2025 (Best Poster Award)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22850 2025-12-01 cs.CV 57%

Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding

解决证据稀疏性:面向长文档理解的代理情境工程

Keliang Liu, Zizhi Chen, Mingcheng Li, Jingqun Tang, Dingkang Yang, Lihua Zhang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 SLEUTH通过多代理框架解决长文档理解中的证据稀疏问题,提升多模态文档处理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22470 2025-12-01 cs.CV 57%

Hybrid, Unified and Iterative: A Novel Framework for Text-based Person Anomaly Retrieval

混合、统一和迭代:一种基于文本的人员异常检索新框架

Tien-Huy Nguyen, Huu-Loc Tran, Huu-Phong Phan-Nguyen, Quang-Vinh Dinh

机构 * University of Information Technology Vietnam National University(信息技术大学越南国家大学) AI VIETNAM Lab(AI越南实验室)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

AI总结 本文提出了一种混合、统一和迭代的新型框架,通过结合局部-全局视角和统一图像-文本模型,提升基于文本的人员异常检索性能。

Comments Accepted on World Wide Web 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22253 2025-12-01 cs.IR cs.CV 57%

UNION: A Lightweight Target Representation for Efficient Zero-Shot Image-Guided Retrieval with Optional Textual Queries

UNION: 一种轻量级的目标表示用于高效零样本图像引导检索与可选文本查询

Hoang-Bao Le, Allie Tran, Binh T. Nguyen, Liting Zhou, Cathal Gurrin

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 UNION通过融合图像嵌入与空文本提示,实现高效零样本图像引导检索,仅需5,000训练样本即在多个基准测试中超越传统基线。

Comments Accepted at ICDM - MMSR Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏