arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-25 至 2025-09-25 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 13 篇

2509.19875 2025-09-25 cs.CV cs.AI 84%

Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection

Yunqing Hu, Zheming Yang, Chang Zhao, Wen Ji

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of AI for Industries(工业人工智能研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20022 2025-09-25 cs.CV 83%

PS3: A Multimodal Transformer Integrating Pathology Reports with Histology Images and Biological Pathways for Cancer Survival Prediction

Manahil Raza, Ayesha Azam, Talha Qaiser, Nasir Rajpoot

机构 * University of Warwick, UK(沃里克大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at ICCV 2025. Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19628 2025-09-25 cs.CE cs.CL q-fin.CP 83%

Multimodal Language Models with Modality-Specific Experts for Financial Forecasting from Interleaved Sequences of Text and Time Series

Ross Koval, Nicholas Andrews, Xifeng Yan

机构 * University of California, Santa Barbara(加州大学圣巴巴拉分校) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20225 2025-09-25 cs.IR cs.AI 79%

Multimodal Representation-disentangled Information Bottleneck for Multimodal Recommendation

Hui Wang, Jinghui Qin, Wushao Wen, Qingling Li, Shanshan Zhong, Zhongzhan Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19719 2025-09-25 cs.CV 79%

Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation

Bo Yu, Jianhua Yang, Zetao Du, Yan Huang, Chenglong Li, Liang Wang

机构 * School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院) NLPR, MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Artificial Intelligence, Anhui University(安徽大学人工智能学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20306 2025-09-25 cs.AI eess.IV q-bio.QM 79%

Multi-Modal Artificial Intelligence of Embryo Grading and Pregnancy Prediction in Assisted Reproductive Technology: A Review

Xueqiang Ouyang, Jia Wei

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10679 2025-09-25 cs.CV 70%

SMLNet: A SPD Manifold Learning Network for Infrared and Visible Image Fusion

Huan Kang, Hui Li, Tianyang Xu, Xiao-Jun Wu, Rui Wang, Chunyang Cheng, Josef Kittler

机构 * School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院) Jiangnan University(江南大学) Centre for Vision, Speech and Signal Processing(视觉、语音与信号处理中心) University of Surrey(Surrey大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 23 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20240 2025-09-25 cs.LG cs.AI 57%

A HyperGraphMamba-Based Multichannel Adaptive Model for ncRNA Classification

Xin An, Ruijie Li, Qiao Ning, Hui Li, Qian Ma, Shikai Guo

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments 9 pages, 17 figures (including subfigures), 1 table. Xin An and Ruijie Li contributed equally to this work and should be considered co-first authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13739 2025-09-25 cs.CV 57%

Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector

Yiming Cao, Yanjie Li, Kaisheng Liang, Bin Xiao

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19733 2025-09-25 cs.CV 57%

Robust RGB-T Tracking via Learnable Visual Fourier Prompt Fine-tuning and Modality Fusion Prompt Generation

Hongtao Yang, Bineng Zhong, Qihua Liang, Zhiruo Zhu, Yaozong Zheng, Ning Li

机构 * Key Laboratory of Education Blockchain and Intelligent Technology, Ministry of Education, Guangxi Normal University(教育区块链与智能技术重点实验室,教育部,广西师范大学) Guangxi Key Lab of Multi-Source Information Mining and Security, Guangxi Normal University(多源信息挖掘与安全广西重点实验室,广西师范大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by TMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16549 2025-09-25 cs.CV 57%

Efficient Rectified Flow for Image Fusion

Zirui Wang, Jiayi Zhang, Tianwei Guan, Yuhan Zhou, Xingyuan Li, Minjing Dong, Jinyuan Liu

机构 * City University of Hong Kong(香港城市大学) Dalian University of Technology(大连理工大学) Chinese University of Hong Kong(香港中文大学) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19860 2025-09-25 cs.CV cs.LG 57%

SpaRC: Sparse Radar-Camera Fusion for 3D Object Detection

Philipp Wolters, Johannes Gilg, Torben Teepe, Fabian Herzog, Felix Fent, Gerhard Rigoll

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 18 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19521 2025-09-25 cs.RO 50%

A Bimanual Gesture Interface for ROS-Based Mobile Manipulators Using TinyML and Sensor Fusion

Najeeb Ahmed Bhuiyan, M. Nasimul Huq, Sakib H. Chowdhury, Rahul Mangharam

机构 * †‡Department of Mechatronics Engineering, Rajshahi University of Engineering \& Technology, Kazla, Rajshahi-6204, Bangladesh §Department of Electrical \& Systems Engineering, School of Engineering \& Applied Science, University of Pennsylvania, Philadelphia, PA 19104, United States Email: , †, ‡, §

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments 12 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏