arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-04 至 2025-11-04 共收录 110 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 16 篇

2511.00859 2025-11-04 cs.CV 79%

Layer-Wise Modality Decomposition for Interpretable Multimodal Sensor Fusion

Jaehyun Park, Konyul Park, Daehun Kim, Junseo Park, Jun Won Choi

机构 * Seoul National University(首尔国立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19769 2025-11-04 cs.CV 79%

AIM: Adaptive Intra-Network Modulation for Balanced Multimodal Learning

Shu Shen, C. L. Philip Chen, Tong Zhang

机构 * Guangdong Provincial Key Laboratory of Computational AI Models and Cognitive Intelligence(广东省计算人工智能模型与认知智能重点实验室) School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) Pazhou Lab(琶洲实验室) Engineering Research Center of the Ministry of Education on Health Intelligent Perception and Paralleled Digital-Human(教育部健康智能感知与并行数字人工程研究中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 13pages,7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01357 2025-11-04 cs.CV cs.AI 79%

CMI-MTL: Cross-Mamba interaction based multi-task learning for medical visual question answering

Qiangguo Jin, Xianyao Zheng, Hui Cui, Changming Sun, Yuqi Fang, Cong Cong, Ran Su, Leyi Wei, Ping Xuan, Junbo Wang

机构 * School of Software, Northwestern Polytechnical University, Shaanxi, China(西北工业大学软件学院) Yangtze River Delta Research Institute of Northwestern Polytechnical University, Taicang, China(西北工业大学长江三角研究 institute) Department of Computer Science and Information Technology, La Trobe University, Melbourne, Australia(拉筹伯大学计算机科学与信息技术系) CSIRO Data61, Sydney, Australia(CSIRO Data61) School of Intelligence Science and Technology, Nanjing University, Suzhou, China(南京大学智能科学与技术学院) Australian Institute of Health Innovation (AIHI), Macquarie University, Australia(麦考瑞大学健康创新研究所) School of Computer Software, College of Intelligence and Computing, Tianjin University, Tianjin, China(天津大学计算机软件学院) Centre for Artificial Intelligence driven Drug Discovery, Faculty of Applied Science, Macao Polytechnic University, Macao Special Administrative Region of China(澳门理工学院人工智能驱动药物发现中心) Department of Computer Science, School of Engineering, Shantou University, Guangdong, China(汕头大学计算机科学系)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments The paper has been accepted by the 33rd Pacific Conference on Computer Graphics and Applications (Pacific Graphics 2025)

Journal ref PG2025 Conference Papers, Posters, and Demos, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00949 2025-11-04 cs.LG 78%

Motion-Robust Multimodal Fusion of PPG and Accelerometer Signals for Three-Class Heart Rhythm Classification

Yangyang Zhao, Matti Kaisti, Olli Lahdenoja, Tero Koivisto

机构 * Department of Computing, Faculty of Technology, University of Turku(图波大学计算系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted for publication in the Companion of the 2025 ACM International Joint Conference on Pervasive and Ubiquitous Computing and the 2025 International Symposium on Wearable Computers (UbiComp/ISWC 2025 Companion). 5 pages, 3 figures. Author's accepted manuscript (AAM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00509 2025-11-04 cs.AI cs.CR 77%

Reimagining Safety Alignment with An Image

Yifan Xia, Guorui Chen, Wenqian Yu, Zhijiang Li, Philip Torr, Jindong Gu

机构 * School of Information Management, Wuhan University, Wuhan, China(武汉大学信息管理学院) Torr Vision Group, University of Oxford, Oxford, United Kingdom(牛津大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01463 2025-11-04 cs.CV cs.AI cs.GR 73%

HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA

Lei Hu, Yongjing Ye, Shihong Xia

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 5figures. The Thirty-Ninth Annual Conference on Neural Information Processing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01284 2025-11-04 cs.CV cs.AI 73%

Adaptation of Foundation Models for Medical Image Analysis: Strategies, Challenges, and Future Directions

Karma Phuntsho, Abdullah, Kyungmi Lee, Ickjai Lee, Euijoon Ahn

机构 * James Cook University(詹姆斯库克大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01082 2025-11-04 cs.CV cs.AI cs.LG 73%

GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction

Narges Ghasemi, Amir Ziashahabi, Salman Avestimehr, Cyrus Shahabi

机构 * of Computer Science, University of Southern California, Los Angeles, CA, USA Computer Engineering, University of Southern California, Los Angeles, CA, USA

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted to IEEE International Conference on Data Mining (ICDM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00218 2025-11-04 cs.CV cs.AI 62%

DM-QPMNET: Dual-modality fusion network for cell segmentation in quantitative phase microscopy

Rajatsubhra Chakraborty, Ana Espinosa-Momox, Riley Haskin, Depeng Xu, Rosario Porras-Aguilar

机构 * College of Computing and Informatics, University of North Carolina at Charlotte, NC, USA(计算与信息学院,北卡罗来纳大学夏洛特分校) Department of Physics and Optical Science, University of North Carolina at Charlotte, NC, USA(物理与光学科学系,北卡罗来纳大学夏洛特分校) Center for TAIMing AI, University of North Carolina at Charlotte, NC, USA(TAIMing AI中心,北卡罗来纳大学夏洛特分校)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26466 2025-11-04 cs.CV cs.LG 57%

Representation-Level Counterfactual Calibration for Debiased Zero-Shot Recognition

Pei Peng, MingKun Xie, Hang Hao, Tong Jin, ShengJun Huang

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19739 2025-11-04 cs.CV 57%

FUSE: Label-Free Image-Event Joint Monocular Depth Estimation via Frequency-Decoupled Alignment and Degradation-Robust Fusion

Pihai Sun, Junjun Jiang, Yuanqi Yao, Youyu Chen, Wenbo Zhao, Kui Jiang, Xianming Liu

机构 * Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Zhengzhou Research Institute, Harbin Institute of Technology(郑州研究院,哈尔滨工业大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments [IROS 2025, camera ready version]: 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 9 篇

2505.23118 2025-11-04 cs.CL cs.AI 81%

Elicit and Enhance: Advancing Multimodal Reasoning in Medical Scenarios

Zhongzhen Huang, Linjie Mu, Yakun Zhu, Xiangyu Zhao, Shaoting Zhang, Xiaofan Zhang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00997 2025-11-04 cs.CV 79%

MID: A Self-supervised Multimodal Iterative Denoising Framework

Chang Nie, Tianchen Deng, Zhe Liu, Hesheng Wang

机构 * School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(自动化与智能感知学院,上海交通大学) Key Laboratory of System Control and Information Processing, Ministry of Education of China(系统控制与信息处理重点实验室,中华人民共和国教育部)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11807 2025-11-04 cs.CL 79%

Are Multimodal Large Language Models Pragmatically Competent Listeners in Simple Reference Resolution Tasks?

Simeon Junker, Manar Ali, Larissa Koch, Sina Zarrieß, Hendrik Buschmeier

机构 * Bielefeld University(比勒菲尔德大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

Comments To appear in ACL Findings 2025

Journal ref Findings of the Association for Computational Linguistics: ACL 2025, pp. 24101-24109

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25381 2025-11-04 cs.HC 78%

CGM-Led Multimodal Tracking with Chatbot Support: An Autoethnography in Sub-Health

Dongyijie Primo Pan, Lan Luo, Yike Wang, Pan Hui

专题命中 其他多模态 :multimodal(title,abstract)

Comments International Conference on Human-Engaged Computing (ICHEC 2025), Singapore

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00411 2025-11-04 cs.LG cs.AI cs.CV 62%

Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided Sampling

Zenghao Niu, Weicheng Xie, Siyang Song, Zitong Yu, Feng Liu, Linlin Shen

机构 * School of Computer Science & Software Engineering, Shenzhen University, China(深圳大学计算机科学与软件工程学院) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ), Shenzhen, China(广东省人工智能与数字经济发展实验室(深圳)) Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University, China(广东省智能信息处理省级重点实验室) School of Computer Science, University of Exeter, U.K.(埃克塞特大学计算机科学学院) Department of Computing and Information Technology, Great Bay University, China(大鹏大学计算与信息技术系) Computer Vision Institute, School of Artificial Intelligence, Shenzhen University, China(人工智能学院计算机视觉研究所)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments accepted by iccv 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01379 2025-11-04 cs.RO 50%

CM-LIUW-Odometry: Robust and High-Precision LiDAR-Inertial-UWB-Wheel Odometry for Extreme Degradation Coal Mine Tunnels

Kun Hu, Menggang Li, Zhiwen Jin, Chaoquan Tang, Eryi Hu, Gongbo Zhou

机构 * School of Mechatronic Engineering, China University of Mining and Technology(机械电子工程学院,中国矿业大学)

专题命中 其他多模态 :multimodal(abstract)

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19660 2025-11-04 cs.ET q-bio.BM 50%

Machine Olfaction and Embedded AI Are Shaping the New Global Sensing Industry

Andreas Mershin, Nikolas Stefanou, Adan Rotteveel, Matthew Kung, George Kung, Alexandru Dan, Howard Kivell, Zoia Okulova, Zoi Kountouri, Paul Pu Liang

专题命中 其他多模态 :multimodal(abstract)

Comments 23 pages, 116 citations, combination tech review/industry roadmap/white paper on the rise of machine olfaction as an essential AI modality

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13066 2025-11-04 astro-ph.HE 50%

Double-Peaked Optical Afterglow in GRB 110213A Inferring a Magnetized Thick Shell Ejecta

Yo Kusafuka, Kaori Obayashi, Katsuaki Asano, Ryo Yamazaki

专题命中 其他多模态 :multimodal(abstract)

Comments 9 pages, 5 figures, accepted for publication in MNRAS, Magglow is available from https://github.com/yo3-sun/Magglow

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00309 2025-11-04 math.OC 50%

Transit-MP: Transit-Prioritized Max-Pressure Control in Sparse Connected Vehicle Environments

Chaopeng Tan, Hao Liu, Dingshan Sun, Marco Rinaldi, Hans van Lint

专题命中 其他多模态 :multi-modal(abstract)

Comments 36 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏