arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-05 至 2025-08-05 共收录 21 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 21 篇

2508.02479 2025-08-05 cs.CV 85%

Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding

Xinquan Yu, Wei Lu, Xiangyang Luo

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02525 2025-08-05 cs.AI 83%

Accurate and Interpretable Postmenstrual Age Prediction via Multimodal Large Language Model

Qifan Chen, Jin Cui, Cindy Duan, Yushuo Han, Yifei Shi

机构 * King’s College London(伦敦国王学院) Imperial College London(帝国理工学院) Columbia University(哥伦比亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Submitted to the NeurIPS 2025 Workshop GenAI4Health. Conference website: https://aihealth.ischool.utexas.edu/GenAI4HealthNeurips2025/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01644 2025-08-05 cs.MM cs.AI cs.CV cs.SD eess.AS 83%

DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition

Peiyuan Jiang, Yao Liu, Qiao Liu, Zongshun Zhang, Jiaye Yang, Lu Liu, Daibing Yao

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Published in ACM Multimedia 2025. 10 pages, 4 figures

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01805 2025-08-05 cs.NI 82%

M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks

Yongjie Zeng, Hongyang Du

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00926 2025-08-05 cs.LG 82%

Hybrid Hypergraph Networks for Multimodal Sequence Data Classification

Feng Xu, Hui Wang, Yuting Huang, Danwei Zhang, Zizhu Fan

机构 * Feng Xu 1,2(作者1单位) Hui Wang 1(作者1单位) Yuting Huang 3(作者3单位) Danwei Zhang 4(作者4单位) Zizhu Fan 5(作者5单位)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01525 2025-08-05 cs.CV cs.AI 81%

MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection

Kuo Shi, Jie Lu, Shanshan Ye, Guangquan Zhang, Zhen Fang

机构 * University of Technology Sydney(悉尼技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01316 2025-08-05 cs.CV cs.HC 79%

Multimodal Attention-Aware Fusion for Diagnosing Distal Myopathy: Evaluating Model Interpretability and Clinician Trust

Mohsen Abbaspour Onari, Lucie Charlotte Magister, Yaoxin Wu, Amalia Lupi, Dario Creazzo, Mattia Tordin, Luigi Di Donatantonio, Emilio Quaia, Chao Zhang, Isel Grau, Marco S. Nobile, Yingqian Zhang, Pietro Liò

机构 * Information Systems Group, Eindhoven University of Technology, The Netherlands(埃因霍温技术大学信息系统组) Eindhoven Artificial Intelligence Systems Institute, The Netherlands(埃因霍温人工智能系统研究所) Department of Computer Science and Technology, University of Cambridge, United Kingdom(剑桥大学计算机科学与技术系) Department of Medicine - DIMED, Padua University Hospital, Italy(帕多瓦大学医院医学部-DIMED) Department of Environmental Sciences, Informatics, and Statistics, Ca’ Foscari University of Venice, Italy(威尼斯卡弗里大学环境科学、信息学与统计学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00963 2025-08-05 cs.LG cs.AI 79%

Rethinking Multimodality: Optimizing Multimodal Deep Learning for Biomedical Signal Classification

Timothy Oladunni, Alex Wong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00830 2025-08-05 cs.CV 79%

Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

Yang Cao, Yihan Zeng, Hang Xu, Dan Xu

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科技大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Code Page: NeurIPS2023" target="_blank" rel="noopener">https://github.com/yangcaoai/CoDA_NeurIPS2023 This paper is accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.07348 2025-08-05 cs.CV 79%

Juggling With Representations: On the Information Transfer Between Imagery, Point Clouds, and Meshes for Multi-Modal Semantics

Dominik Laupheimer, Norbert Haala

机构 * Institute for Photogrammetry, University of Stuttgart, Germany(摄影测量研究所,斯图加特大学,德国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02409 2025-08-05 cs.CV cs.AI 76%

Hydra: Accurate Multi-Modal Leaf Wetness Sensing with mm-Wave and Camera Fusion

Yimeng Liu, Maolin Gan, Huaili Zeng, Li Liu, Younsuk Dong, Zhichao Cao

机构 * Michigan State University(密歇根州立大学) Tsinghua University(清华大学)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

Comments In Proceedings of ACM MobiCom (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01150 2025-08-05 cs.CV 70%

OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level Understanding

Dianyi Yang, Xihan Wang, Yu Gao, Shiyang Liu, Bohan Ren, Yufeng Yue, Yi Yang

机构 * School of Automation, Beijing Institute of Technology, Beijing, China(自动化学院,北京理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments IROS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02331 2025-08-05 cs.CV cs.SD 70%

VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection

Hao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu, Jia Li, Meng Wang, Richang Hong

机构 * Hefei University of Technology(合肥工业大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Source code and pre-trained models will be available at https://github.com/MSA-LMC/VAEmo

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19903 2025-08-05 cs.CV 70%

Scaling Vision Pre-Training to 4K Resolution

Baifeng Shi, Boyi Li, Han Cai, Yao Lu, Sifei Liu, Marco Pavone, Jan Kautz, Song Han, Trevor Darrell, Pavlo Molchanov, Hongxu Yin

机构 * UC Berkeley(加州大学伯克利分校) NVIDIA(英伟达)

专题命中 多模态训练与对齐 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments CVPR 2025. Project Page: https://nvlabs.github.io/PS3

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01211 2025-08-05 cs.LG 67%

Multi-Operator Few-Shot Learning for Generalization Across PDE Families

Yile Li, Shandian Zhe

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03257 2025-08-05 cs.CV 57%

LACONIC: A 3D Layout Adapter for Controllable Image Creation

Léopold Maillard, Tom Durand, Adrien Ramanana Rahary, Maks Ovsjanikov

机构 * LIX, École Polytechnique, IP Paris(巴黎高等理工学院LIX研究所) Dassault Systèmes(达索系统)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02423 2025-08-05 q-bio.TO 50%

Evolutionary Paradigms in Histopathology Serial Sections technology

Zhenfeng Zhuang, Min Cen, Lei Jiang, Qiong Peng, Yihuang Hu, Hong-Yu Zhou, Liansheng Wang

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12024 2025-08-05 stat.ML cs.LG hep-ex hep-ph 50%

Comparison of Affine and Rational Quadratic Spline Coupling and Autoregressive Flows through Robust Statistical Tests

Andrea Coccaro, Marco Letizia, Humberto Reyes-Gonzalez, Riccardo Torre

机构 * INFN(意大利国家核物理研究所) MaLGa - DIBRIS(马尔盖-迪布里兹) University of Genova(热那亚大学) Department of Physics(物理系) RWTH Aachen(亚琛工业大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments v3: published version; 25 pages, 3 figures, 3 tables

Journal ref Symmetry 2024, 16(8), 942

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01599 2025-08-05 cs.OH 50%

Lessons Learned from the Real-World Deployment of Multi-Sensor Fusion for Proactive Work Zone Safety Application

Minhaj Uddin Ahmad, Sagar Dasgupta, Mizanur Rahman, Sakib Khan, Md Wasiul Haque, Suhala Rabab Saba, David Bodoh, Nathan Huynh, Li Zhao, Eren Erman Ozguven

专题命中 多模态训练与对齐 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01375 2025-08-05 cs.IR 50%

SaviorRec: Semantic-Behavior Alignment for Cold-Start Recommendation

Yining Yao, Ziwei Li, Shuwen Xiao, Boya Du, Jialin Zhu, Junjun Zheng, Xiangheng Kong, Yuning Jiang

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00862 2025-08-05 cs.DL 50%

A survey on proximity monitoring and warning in construction

Yuexiong Ding, Qiong Liu, Ankang Ji, Xiaowei Luo, Wen Yi, Albert P. C. Chan

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏