arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-17 至 2025-09-17 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 6 篇

2509.12600 2025-09-17 cs.LG cs.AI q-bio.QM 88%

A Multimodal Foundation Model to Enhance Generalizability and Data Efficiency for Pan-cancer Prognosis Prediction

Huajun Zhou, Fengtao Zhou, Jiabo Ma, Yingxue Xu, Xi Wang, Xiuming Zhang, Li Liang, Zhenhui Li, Hao Chen

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Hong Kong University of Science and Technology(香港科学与技术大学) Department of Pathology(病理学系) School of Medicine(医学院) Zhejiang University(浙江大学) Nanfang Hospital and School of Basic Medical Sciences(南方医科大学基础医学系) Southern Medical University(南方医学院) Guangdong Provincial Key Laboratory of Molecular Tumor Pathology(广东省分子肿瘤病理重点实验室) Jinfeng Laboratory(金凤实验室) Department of Radiology(放射科) Division of Life Science(生命科学系) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute(香港科技大学深圳-香港协同创新研究院) State Key Laboratory of Nervous System Disorders(神经系统疾病国家重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.AI

Comments 27 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12653 2025-09-17 cs.CV cs.AI 84%

Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations

Jinjie Shen, Yaxiong Wang, Lechao Cheng, Nan Pu, Zhun Zhong

机构 * Hefei University of Technology(合肥工业大学) University of Trento(特伦托大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01275 2025-09-17 cs.AI 83%

Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D

Artemis Panagopoulou, Le Xue, Honglu Zhou, silvio savarese, Ran Xu, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, Juan Carlos Niebles

机构 * Salesforce AI Reseach(Salesforce人工智能研究院) University of Pennsylvania(宾夕法尼亚大学)

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12994 2025-09-17 cs.CL 70%

SitLLM: Large Language Models for Sitting Posture Health Understanding via Pressure Sensor Data

Jian Gao, Fufangchen Zhao, Yiyang Zhang, Danfeng Yan

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13175 2025-09-17 cs.CV 57%

More performant and scalable: Rethinking contrastive vision-language pre-training of radiology in the LLM era

Yingtai Li, Haoran Lai, Xiaoqian Zhou, Shuai Ming, Wenxin Ma, Wei Wei, Shaohua Kevin Zhou

机构 * School of Biomedical Engineering, Division of Life Sciences Medicine, University of Science Technology of China (USTC), Hefei Anhui, 230026, China Center for Medical Imaging, Robotics, Analytic Computing \& Learning (MIRACLE), Suzhou Institute for Advance Research, USTC, Suzhou Jiangsu, 215123, China The First Affiliated Hospital of USTC, Division of Life Sciences Medicine, USTC, Hefei Anhui, 230001, China Jiangsu Provincial Key Laboratory of Multimodal Digital Twin Technology, Suzhou Jiangsu, 215123, China State Key Laboratory of Precision

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21875 2025-09-17 cs.AI 57%

Tiny-BioMoE: a Lightweight Embedding Model for Biosignal Analysis

Stefanos Gkikas, Ioannis Kyprakis, Manolis Tsiknakis

机构 * Foundation for Research \& Technology-Hellas Heraklion Greece Foundation for Research \& Technology-Hellas Hellenic Mediterranean University Heraklion Greece Hellenic Mediterranean University

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏