arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2507.11892 2025-07-17 cs.CV cs.AI cs.HC 81%

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition

Yu Liu, Leyuan Qu, Hanlei Shi, Di Gao, Yuhua Zheng, Taihao Li

机构 * Hangzhou Institute for Advanced Study(杭州先进研究院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05285 2025-07-15 cs.CL cs.AI cs.CY cs.IR 81%

Beyond classical and contemporary models: a transformative AI framework for student dropout prediction in distance learning using RAG, Prompt engineering, and Cross-modal fusion

Miloud Mihoubi, Meriem Zerkouk, Belkacem Chikhaoui

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CL、cs.AI

Comments 13 pages, 8 figures, 1 Algorithms, 17th International Conference on Education and New Learning Technologies,: 30 June-2 July, 2025 Location: Palma, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06067 2025-07-09 eess.IV cs.AI cs.CV 81%

Enhancing Synthetic CT from CBCT via Multimodal Fusion and End-To-End Registration

Maximilian Tschuchnig, Lukas Lamminger, Philipp Steininger, Michael Gadermayr

机构 * Salzburg University of Applied Sciences(萨尔茨堡应用科学大学) MedPhoton GmbH(MedPhoton公司) University of Salzburg(萨尔茨堡大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at CAIP 2025. arXiv admin note: substantial text overlap with arXiv:2506.08716

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04397 2025-07-01 cs.CL cs.AI cs.LG 81%

Multimodal Medical Code Tokenizer

Xiaorui Su, Shvat Messica, Yepeng Huang, Ruth Johnson, Lukas Fesser, Shanghua Gao, Faryad Sahneh, Marinka Zitnik

机构 * Department of Biomedical Informatics, Harvard Medical School, Boston, MA, USA(生物医学信息学系,哈佛医学院,波士顿,马萨诸塞州,美国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments ICML'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09062 2025-07-01 cs.CV cs.AI cs.RO 81%

Multimodal Object Detection using Depth and Image Data for Manufacturing Parts

Nazanin Mahjourian, Vinh Nguyen

机构 * Department of Mechanical Engineering - Engineering Mechanics, Michigan Technological University(机械工程系-工程力学系,密歇根技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05319 2025-06-26 cs.CV cs.AI 81%

Robust Multimodal Learning for Ophthalmic Disease Grading via Disentangled Representation

Xinkun Wang, Yifang Wang, Senwei Liang, Feilong Tang, Chengzhi Liu, Ming Hu, Chao Hu, Junjun He, Zongyuan Ge, Imran Razzak

机构 * MBZUAI, United Arab Emirates(MBZUAI,阿联酋) Monash University, Australia(墨尔本大学,澳大利亚) Liverpool University, United Kingdom(利物浦大学,英国) China Unicom (Shanghai) Industrial Internet Co., Ltd., China(中国联合(上海)工业互联网有限公司,中国) Shanghai AI Lab, China(上海人工智能实验室,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 10pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18204 2025-06-25 cs.CV cs.AI 81%

Multimodal Fusion SLAM with Fourier Attention

Youjie Zhou, Guofeng Mei, Yiming Wang, Yi Wan, Fabio Poiesi

机构 * School of Mechanical Engineering, Shandong University(机械工程学院,山东大学) Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in IEEE RAL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18683 2025-06-24 cs.CV cs.AI 81%

SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification

Youcef Sklab, Hanane Ariouat, Eric Chenin, Edi Prifti, Jean-Daniel Zucker

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 25 pages, 9 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14704 2025-06-18 cs.CV cs.MM 81%

Tile Classification Based Viewport Prediction with Multi-modal Fusion Transformer

Zhihao Zhang, Yiwei Chen, Weizhan Zhang, Caixia Yan, Qinghua Zheng, Qi Wang, Wangdu Chen

机构 * Xi'an Jiaotong University(西安交通大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments This paper is accepted by ACM-MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11465 2025-06-16 cs.LG cs.AI cs.CV 81%

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer

Haotian Ni, Yake Wei, Hang Liu, Gong Chen, Chong Peng, Hao Lin, Di Hu

机构 * Gaoling School of Artificial Intelligence Renmin University of China(中国人民大学人工智能学院) Beihang University(北京航空航天大学) Xiamen University(厦门大学) Beijing Key Laboratory of Research on Large Models(北京市大模型研究关键实验室) Engineering Research Center of Next-Generation Intelligent Search(下一代智能搜索工程研究中心) Tencent(腾讯)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07943 2025-06-12 cs.CV cs.AI 81%

Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations

Yizhen Li, Dell Zhang, Xuelong Li, Yiqing Shen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments This work was submitted without the consent of all co-authors. We request withdrawal until all parties agree

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02080 2025-06-12 cs.CV cs.CL cs.LG 81%

EMMA: Efficient Visual Alignment in Multi-Modal LLMs

Sara Ghazanfari, Alexandre Araujo, Prashanth Krishnamurthy, Siddharth Garg, Farshad Khorrami

机构 * Department of Electronic and Computer Engineering, New York University(电子与计算机工程系,纽约大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11494 2025-06-10 cs.CL cs.CV 81%

Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More

Zichen Wen, Yifeng Gao, Shaobo Wang, Junyuan Zhang, Qintong Zhang, Weijia Li, Conghui He, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Sun Yat-sen University(中山大学) Peking University(北京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 15 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04788 2025-06-06 cs.CL cs.AI cs.LG 81%

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques

Jisu An, Junseok Lee, Jeoungeun Lee, Yongseok Son

机构 * Seoul National University(首尔国立大学) University of California San Diego(加州大学圣地亚哥分校) Chung-Ang University(Chung-Ang 大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 18 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18053 2025-06-06 cs.CL cs.CV 81%

DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models

Jianyu Liu, Hangyu Guo, Ranjie Duan, Xingyuan Bu, Yancheng He, Shilong Li, Hui Huang, Jiaheng Liu, Yucheng Wang, Chenchen Jing, Xingwei Qu, Xiao Zhang, Yingshui Tan, Yanan Wu, Jihao Gu, Yangguang Li, Jianke Zhu

机构 * Alibaba Group(阿里巴巴集团) Zhejiang University(浙江大学) M-A-P The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments [NAACL 2025] The first four authors contribute equally, 23 pages, repo at https://github.com/Kizna1ver/DREAM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15358 2025-06-05 cs.CL cs.CV 81%

SemEval-2025 Task 1: AdMIRe -- Advancing Multimodal Idiomaticity Representation

Thomas Pickard, Aline Villavicencio, Maggie Mi, Wei He, Dylan Phelps, Marco Idiart

机构 * University of Sheffield, UK(谢菲尔德大学) University of Exeter, UK(埃克塞特大学) Federal University of Rio Grande do Sul, Brazil(里约格朗德杜斯尔大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Author accepted version; SemEval-2025 proceedings to appear at ACL 2025. This version corrects a typo in the results table

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05237 2025-06-05 cs.CL cs.CV 81%

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Jarvis Guo, Tuney Zheng, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Yizhi Li, Graham Neubig, Wenhu Chen, Xiang Yue

机构 * Carnegie Mellon University(卡内基梅隆大学) M-A-P Nanyang Technological University(南洋理工大学) University of Waterloo(滑铁卢大学) The University of Manchester(曼彻斯特大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16282 2025-06-03 eess.IV cs.AI cs.CV 81%

Brain-Adapter: Enhancing Neurological Disorder Analysis with Adapter-Tuning Multimodal Large Language Models

Jing Zhang, Xiaowei Yu, Yanjun Lyu, Lu Zhang, Tong Chen, Chao Cao, Yan Zhuang, Minheng Chen, Tianming Liu, Dajiang Zhu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22271 2025-05-29 cs.CR cs.AI cs.CL 81%

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models

Yongcan Yu, Yanbo Wang, Ran He, Jian Liang

机构 * NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences(国家工程实验室与人工智能研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18531 2025-05-27 cs.AI cs.CV 81%

Generative RLHF-V: Learning Principles from Multi-modal Human Preference

Jiayi Zhou, Jiaming Ji, Boyuan Chen, Jiapeng Sun, Wenqi Chen, Donghai Hong, Sirui Han, Yike Guo, Yaodong Yang

机构 * Peking University(北京大学) Hong Kong University of Science and Technology(香港科技大学) University College London(伦敦大学学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14499 2025-05-27 cs.CL cs.AI 81%

Enhanced Multimodal Aspect-Based Sentiment Analysis by LLM-Generated Rationales

Jun Cao, Jiyi Li, Ziwei Yang, Renjie Zhou

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 15 pages, 2 figures, 6 tables. Accepted by ICONIP2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15687 2025-05-22 cs.CV cs.AI 81%

Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning

Zhe Xu, Cheng Jin, Yihui Wang, Ziyi Liu, Hao Chen

机构 * Department of Computer Science Engineering(计算机科学与工程系) HKUST(香港科技大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13210 2025-05-20 cs.CL cs.AI 81%

Picturized and Recited with Dialects: A Multimodal Chinese Representation Framework for Sentiment Analysis of Classical Chinese Poetry

Xiaocong Du, Haoyu Pei, Haipeng Zhang

机构 * ShanghaiTech University(上海科技大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11066 2025-05-19 cs.AI cs.MM 81%

A Multi-modal Fusion Network for Terrain Perception Based on Illumination Aware

Rui Wang, Shichun Yang, Yuyi Chen, Zhuoyang Li, Zexiang Tong, Jianyi Xu, Jiayi Lu, Xinjie Feng, Yaoguang Cao

机构 * Department of Transportation Science and Engineering, Beihang University(北京航空航天大学交通运输科学与工程学院) Innovation Center of New Energy Vehicle Digital Supervision Technology and Application for State Market Regulation(国家新能源汽车数字监管技术及应用创新中心) Hangzhou International Innovation Institute, Beihang University(杭州国际创新研究院) State Key Lab of Intelligent Transportation System, Beihang University(智能交通系统国家重点实验室)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19458 2025-05-16 cs.MM cs.CL cs.IR 81%

Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal Perspective

Taoyu Su, Jiawei Sheng, Duohe Ma, Xiaodong Li, Juwei Yue, Mengxiao Song, Yingkai Tang, Tingwen Liu

机构 * Institute of Information Engineering, Chinese Academy of Sciences(信息工程研究所,中国科学院) School of Cyber Security, University of Chinese Academy of Sciences(网络安全学院,中国科学院大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL、cs.MM

Comments Accepted by SIGIR 2025, 11 pages, 10 figures, 4 tables,

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04545 2025-05-09 cs.MM cs.CL 81%

TCAN: Text-oriented Cross Attention Network for Multimodal Sentiment Analysis

Weize Quan, Yunfei Feng, Ming Zhou, Yunzhen Zhao, Tong Wang, Dong-Ming Yan

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) Tencent(腾讯) School of Information Science and Technology, Donghua University(东华大学信息科学与技术学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01615 2025-05-06 cs.CV cs.AI 81%

Multimodal and Multiview Deep Fusion for Autonomous Marine Navigation

Dimitrios Dagdilelis, Panagiotis Grigoriadis, Roberto Galeazzi

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20984 2025-05-02 cs.CL cs.AI 81%

UoR-NCL at SemEval-2025 Task 1: Using Generative LLMs and CLIP Models for Multilingual Multimodal Idiomaticity Representation

Thanet Markchom, Tong Wu, Liting Huang, Huizhi Liang

机构 * Department of Computer Science, University of Reading(阅读大学计算机科学系) School of Computing, Newcastle University(新castle大学计算机学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14684 2025-04-29 eess.IV cs.AI cs.CV 81%

Learning Modality-Aware Representations: Adaptive Group-wise Interaction Network for Multimodal MRI Synthesis

Tao Song, Yicheng Wu, Minhao Hu, Xiangde Luo, Linda Wei, Guotai Wang, Yi Guo, Feng Xu, Shaoting Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17822 2025-04-28 cs.CV cs.AI 81%

A multi-scale vision transformer-based multimodal GeoAI model for mapping Arctic permafrost thaw

Wenwen Li, Chia-Yu Hsu, Sizhe Wang, Zhining Gu, Yili Yang, Brendan M. Rogers, Anna Liljedahl

机构 * School of Geographical Sciences and Urban Planning Arizona State University(地理科学与城市规划学院亚利桑那州立大学) Woodwell Climate Research Center(伍德沃德气候研究中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏