arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-16 至 2025-09-16 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 9 篇

2504.10307 2025-09-16 cs.IR 89%

CROSSAN: Towards Efficient and Effective Adaptation of Multiple Multimodal Foundation Models for Sequential Recommendation

Junchen Fu, Yongxin Ni, Joemon M. Jose, Ioannis Arapakis, Kaiwen Zheng, Youhua Li, Xuri Ge

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06641 2025-09-16 cs.AI cs.LG 83%

CogGuide: Human-Like Guidance for Zero-Shot Omni-Modal Reasoning

Zhou-Peng Shou, Zhi-Qiang You, Fang Wang, Hai-Bo Liu

机构 * NoDesk AI Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :omni-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10986 2025-09-16 cs.CV cs.RO 79%

Long-Tailed 3D Detection via Multi-Modal Fusion

Yechi Ma, Neehar Peri, Achal Dave, Wei Hua, Deva Ramanan, Shu Kong

机构 * Department of Computer Science(计算机科学系) Robotics Institute(机器人研究所) Toyota Research Institute(丰田研究机构) Faculty of Science and Technology(科学与技术学院) Institute of Collaborative Innovation(协同创新研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments The first two authors contributed equally. Project page: https://mayechi.github.io/lt3d-lf-io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11376 2025-09-16 cs.LG cs.AI cs.CE 79%

Intelligent Reservoir Decision Support: An Integrated Framework Combining Large Language Models, Advanced Prompt Engineering, and Multimodal Data Fusion for Real-Time Petroleum Operations

Seyed Kourosh Mahjour, Seyed Saman Mahjour

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11187 2025-09-16 cs.CR 78%

DMLDroid: Deep Multimodal Fusion Framework for Android Malware Detection with Resilience to Code Obfuscation and Adversarial Perturbations

Doan Minh Trung, Tien Duc Anh Hao, Luong Hoang Minh, Nghi Hoang Khoa, Nguyen Tan Cam, Van-Hau Pham, Phan The Duy

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11817 2025-09-16 cs.CV 57%

MAFS: Masked Autoencoder for Infrared-Visible Image Fusion and Semantic Segmentation

Liying Wang, Xiaoli Zhang, Chuanmin Jia, Siwei Ma

机构 * Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(教育部符号计算与知识工程重点实验室,吉林大学) Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) National Engineering Research Center of Visual Technology, School of Computer Science, Peking University(视觉技术国家工程研究中心,北京大学计算机科学学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted by TIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11476 2025-09-16 cs.CV cs.LG 57%

Modality-Aware Infrared and Visible Image Fusion with Target-Aware Supervision

Tianyao Sun, Dawei Xiang, Tianqi Ding, Xiang Fang, Yijiashun Qi, Zunduo Zhao

机构 * Independent researcher(独立研究者) Dept. of Computer Science Baylor University(计算机科学系 巴里尔大学) Dept. of Computer Science Engineering University of Connecticut(计算机科学工程系 佛罗里达大学) Dept. of Computer Science New York University(计算机科学系 新 york 大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by 2025 6th International Conference on Computer Vision and Data Mining (ICCVDM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11102 2025-09-16 cs.CV 57%

Filling the Gaps: A Multitask Hybrid Multiscale Generative Framework for Missing Modality in Remote Sensing Semantic Segmentation

Nhi Kieu, Kien Nguyen, Arnold Wiliem, Clinton Fookes, Sridha Sridharan

机构 * School of Electrical Engineering and Robotics, Queensland University of Technology(电气工程与机器人学学院,昆士兰理工大学) Shield AI

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted to DICTA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09085 2025-09-16 cs.CV 57%

IRDFusion: Iterative Relation-Map Difference guided Feature Fusion for Multispectral Object Detection

Jifeng Shen, Haibo Zhan, Xin Zuo, Heng Fan, Xiaohui Yuan, Jun Li, Wankou Yang

机构 * School of Electrical and Information Engineering, Jiangsu University, Zhenjiang, 212013, China(江苏大学电气与信息工程学院) School of Computer Science and Engineering, Jiangsu University of Science and Technology, Zhenjiang, 212003, China(江苏科技大学计算机科学与工程学院) School of Automation, Southeast University, Nanjing, 210096, China(东南大学自动化学院) Department of Computer Science and Engineering, University of North Texas, Denton, TX 76207, USA(北卡罗来纳州立大学计算机科学与工程系) School of Computing, Nanjing Normal University, Nanjing, 210046, Jiangsu(南京师范大学计算机学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 31 pages,6 figures, submitted on 3 Sep,2025

详情

展开后加载摘要…

URL PDF HTML 收藏