arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-26 至 2025-08-26 共收录 15 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 15 篇

2410.05849 2025-08-26 cs.CV 83%

ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt

Fanhu Zeng, Fei Zhu, Haiyang Guo, Xu-Yao Zhang, Cheng-Lin Liu

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(多模态人工智能系统国家重点实验室) School of Artificial Intelligence, UCAS(人工智能学院) Centre for Artificial Intelligence and Robotics, HKISI-CAS(人工智能与机器人中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16744 2025-08-26 cs.LG cs.CL cs.CV 81%

Hyperbolic Multimodal Representation Learning for Biological Taxonomies

ZeMing Gong, Chuanqi Tang, Xiaoliang Huo, Nicholas Pellegrino, Austin T. Wang, Graham W. Taylor, Angel X. Chang, Scott C. Lowe, Joakim Bruslund Haurum

机构 * Simon Fraser University(西蒙弗雷泽大学) University of Waterloo(滑铁卢大学) Vector Institute(向量研究所) University of Guelph(圭尔夫大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所) Aalborg University(奥胡斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18050 2025-08-26 cs.CV 79%

ArgusCogito: Chain-of-Thought for Cross-Modal Synergy and Omnidirectional Reasoning in Camouflaged Object Segmentation

Jianwen Tan, Huiyao Zhang, Rui Xiong, Han Zhou, Hongfei Wang, Ye Li

机构 * University of Chinese Academy of Sciences(中国科学院大学) Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间利用技术与工程中心)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17478 2025-08-26 cs.CV 79%

GraphMMP: A Graph Neural Network Model with Mutual Information and Global Fusion for Multimodal Medical Prognosis

Xuhao Shan, Ruiquan Ge, Jikui Liu, Linglong Wu, Chi Zhang, Siqi Liu, Wenjian Qin, Wenwen Min, Ahmed Elazab, Changmiao Wang

机构 * Hangzhou Dianzi University(杭州电子科技大学) Shenzhen Polytechnic University(深圳职业技术大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Yunnan University(云南大学) Shenzhen University(深圳大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17213 2025-08-26 cs.CV 79%

Multi-modal Knowledge Decomposition based Online Distillation for Biomarker Prediction in Breast Cancer Histopathology

Qibin Zhang, Xinyu Hao, Qiao Chen, Rui Xu, Fengyu Cong, Cheng Lu, Hongming Xu

机构 * School of Biomedical Engineering, Faulty of Medicine, Dalian University of Technology, Dalian, China(生物医学工程学院) Faculty of Information Technology, University of Jyvaskyla, Jyvaskyla, Finland(信息科技学院) School of Software Technology, Dalian University of Technology, Dalian, China(软件技术学院) Department of Radiology, Guangdong Provincial People’s Hospital, Southern Medical University, Guangzhou, China(放射科)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17640 2025-08-26 eess.SP 78%

Multimodal Radio and Vision Fusion for Robust Localization in Urban V2I Communications

Can Zheng, Jiguang He, Chung G. Kang, Guofa Cai, Henk Wymeersch

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 6 pages, 6 figures, submitted to conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17524 2025-08-26 cs.CV cs.AI 73%

OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation

Xingxin He, Aurora Rofena, Ruimin Feng, Haozhe Liao, Zhaoye Zhou, Albert Jang, Fang Liu

机构 * Athinoula A. Martinos Center for Biomedical Imaging(阿提诺拉A.马丁诺斯生物医学成像中心) Harvard Medical School(哈佛医学院) Massachusetts General Hospital(麻省总医院) University Campus Bio-Medico of Rome(罗马生物医学大学校园)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17638 2025-08-26 cs.CV cs.CL 62%

Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning

Xinyu Wei, Guoli Yang, Jialu Zhou, Mingyue Yang, Leqian Li, Kedi Zhang, Chunping Qiu

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15123 2025-08-26 cs.CV cs.AI 62%

Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding

Ta Duc Huy, Duy Anh Huynh, Yutong Xie, Yuankai Qi, Qi Chen, Phi Le Nguyen, Sen Kim Tran, Son Lam Phung, Anton van den Hengel, Zhibin Liao, Minh-Son To, Johan W. Verjans, Vu Minh Hieu Phan

机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Macquarie University(麦考瑞大学) Hanoi University of Science and Technology(河内科学技术大学) University of Wollongong(沃林根大学) Flinders University(弗林德斯大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at ICCV 2025 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17619 2025-08-26 cs.CV 57%

Improving Interpretability in Alzheimer's Prediction via Joint Learning of ADAS-Cog Scores

Nur Amirah Abd Hamid, Mohd Shahrizal Rusli, Muhammad Thaqif Iman Mohd Taufek, Mohd Ibrahim Shapiai, Daphne Teck Ching Lai

机构 * School of Digital Science(数字科学学院) Universiti Brunei Darussalam(布鲁尼岛大学) Faculty of Artificial Intelligence(人工智能学院) Universiti Teknologi Malaysia(马来西亚理工大学) Malaysia-Japan International Institute of Technology(马来西亚-日本国际理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17283 2025-08-26 cs.CV cs.LG 57%

Quickly Tuning Foundation Models for Image Segmentation

Breenda Das, Lennart Purucker, Timur Carstensen, Frank Hutter

机构 * University of Freiburg(弗赖堡大学) ELLIS Institute Tübingen(图宾根ELLIS研究所) Prior Labs(Prior实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted as a short paper at the non-archival content track of AutoML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16974 2025-08-26 cs.CV 57%

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding

Leilei Guo, Antonio Carlos Rivera, Peiyu Tang, Haoxuan Ren, Zheyu Song

机构 * Zhongkai University of Agriculture and Engineering(仲恺农业工程大学) EDP University of Puerto Rico: San Sebastian(波多黎各圣塞巴斯蒂安EDP大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03430 2025-08-26 cs.LG cs.AI 57%

Multi-Level Fusion Graph Neural Network for Molecule Property Prediction

XiaYu Liu, Chao Fan, Yang Liu, Hou-biao Li

机构 * School of Mathematical Sciences, University of Electronic Science and Technology of China(电子科技大学数学科学学院) College of Management Science, Chengdu University of Technology(成都理工大学管理科学学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments 42 pages, 11 figures, 6 tables

Journal ref Journal of Chemical Information and Modeling,2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16579 2025-08-26 cs.CV 57%

Towards High-Precision Depth Sensing via Monocular-Aided iToF and RGB Integration

Yansong Du, Yutong Deng, Yuting Zhou, Feiyu Jiao, Jian Song, Xun Guan

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18058 2025-08-26 q-bio.QM 50%

Comprehensively stratifying MCIs into distinct risk subtypes based on brain imaging genetics fusion learning

Muheng Shang, Jin Zhang, Junwei Han, Lei Du

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏