arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-29 至 2025-07-29 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 16 篇

2507.20620 2025-07-29 cs.AI cs.CV 84%

Complementarity-driven Representation Learning for Multi-modal Knowledge Graph Completion

Lijian Li

机构 * Department of Computer and Information Science, University of Macau(计算机与信息科学系,澳门大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03726 2025-07-29 cs.CV cs.CL 84%

Otter: A Multi-Modal Model with In-Context Instruction Tuning

Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Joshua Adrian Cahyono, Jingkang Yang, Ziwei Liu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19839 2025-07-29 cs.LG cs.CV 83%

GNSP: Gradient Null Space Projection for Preserving Cross-Modal Alignment in VLMs Continual Learning

Tiantian Peng, Yuyang Liu, Shuo Yang, Qiuhe Hong, YongHong Tian

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20749 2025-07-29 cs.CL cs.CV 81%

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study

Yiran Huang, Lukas Thede, Massimiliano Mancini, Wenjia Xu, Zeynep Akata

机构 * Technical University of Munich, Germany(慕尼黑技术大学,德国) Helmholtz Munich, Munich Center for Machine Learning, Germany(海德堡慕尼黑,慕尼黑机器学习中心,德国) University of Tübingen, Tübingen AI Center, Germany(图宾根大学,图宾根人工智能中心,德国) University of Trento, Italy(特伦托大学,意大利) Beijing University of Posts and Telecommunications, China(北京邮电大学,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at GCPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06821 2025-07-29 cs.LG cs.AI cs.MM 81%

HeLo: Heterogeneous Multi-Modal Fusion with Label Correlation for Emotion Distribution Learning

Chuhang Zheng, Chunwei Tian, Jie Wen, Daoqiang Zhang, Qi Zhu

机构 * College of Artificial Intelligence(人工智能学院) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) School of Computer Science and Technology(计算机科学与技术学院) Harbin Institute of Technology(哈尔滨工业大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳分校)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20714 2025-07-29 cs.LG cs.AI q-bio.QM stat.AP 79%

Prostate Cancer Classification Using Multimodal Feature Fusion and Explainable AI

Asma Sadia Khan, Fariba Tasnia Khan, Tanjim Mahmud, Salman Karim Khan, Rishita Chakma, Nahed Sharmen, Mohammad Shahadat Hossain, Karl Andersson

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14683 2025-07-29 cs.CV 79%

Emerging Properties in Unified Multimodal Pretraining

Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, Guang Shi, Haoqi Fan

机构 * ByteDance Seed(字节跳动种子) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Monash University(墨尔本大学) Hong Kong University of Science and Technology(香港科学与技术大学) UC Santa Cruz(加州大学圣克ruz分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 37 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20518 2025-07-29 cs.CV cs.MM 79%

T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval

Yili Li, Gang Xiong, Gaopeng Gou, Xiangyan Qu, Jiamin Zhuang, Zhen Li, Junzheng Shi

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19679 2025-07-29 cs.CV cs.AI 76%

Efficient Learning for Product Attributes with Compact Multimodal Models

Mandar Kulkarni

机构 * Flipkart Data Science(Flipkart数据科学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10887 2025-07-29 cs.CV cs.AI 73%

Point Cloud Self-supervised Learning via 3D to Multi-view Masked Learner

Zhimin Chen, Xuewei Chen, Xiao Guo, Yingwei Li, Longlong Jing, Liang Yang, Bing Li

机构 * Clemson University(克莱姆森大学) Michigan State University(密歇根州立大学) Johns Hopkins University(约翰霍普金斯大学) The City University of New York(纽约城市大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20842 2025-07-29 cs.CV 57%

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models

Yuchen Liu, Yaoming Wang, Bowen Shi, Xiaopeng Zhang, Wenrui Dai, Chenglin Li, Hongkai Xiong, Qi Tian

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20025 2025-07-29 cs.CV 57%

Region-based Cluster Discrimination for Visual Representation Learning

Yin Xie, Kaicheng Yang, Xiang An, Kun Wu, Yongle Zhao, Weimo Deng, Zimin Ran, Yumeng Wang, Ziyong Feng, Roy Miles, Ismail Elezi, Jiankang Deng

机构 * DeepGlint University of Technology Sydney(悉尼科技大学) Huawei London Research Center(华为伦敦研究中心) Imperial College London(伦敦帝国理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted as a highlight paper at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22373 2025-07-29 cs.LG cs.AI 57%

Analytic Continual Test-Time Adaptation for Multi-Modality Corruption

Yufei Zhang, Yicheng Xu, Hongxin Wei, Zhiping Lin, Xiaofeng Zou, Cen Chen, Huiping Zhuang

机构 * South China University of Technology(华南理工大学) Institute of Science Tokyo(东京科学研究所) Southern University of Science and Technology(南方科技大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19840 2025-07-29 cs.CV cs.AI cs.CL 56%

AutoSign: Direct Pose-to-Text Translation for Continuous Sign Language Recognition

Samuel Ebimobowei Johnny, Blessed Guda, Andrew Blayama Stephen, Assane Gueye

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态训练与对齐 :分类 cs.CV、cs.CL、cs.AI;multimodal(comments)

Comments Paper to appear at the 1st Workshop in Multimodal Sign Language Recognition at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20925 2025-07-29 cs.LG q-bio.QM 50%

Zero-Shot Learning with Subsequence Reordering Pretraining for Compound-Protein Interaction

Hongzhi Zhang, Zhonglie Liu, Kun Meng, Jiameng Chen, Jia Wu, Bo Du, Di Lin, Yan Che, Wenbin Hu

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Engineering Research Center for Big Data Application in Private Health Medicine of Fujian Universities, Putian University(福建省私人健康医学大数据应用工程研究中心,莆田大学) Macquarie University(麦考瑞大学) Wuhan University Shenzhen Research Institute(武汉大学深圳研究院)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20089 2025-07-29 cs.LG stat.ME stat.ML 50%

Meta Fusion: A Unified Framework For Multimodality Fusion with Mutual Learning

Ziyi Liang, Annie Qu, Babak Shahbaba

机构 * Department of Statistics, University of California, Irvine, CA, USA(统计学系,加州大学伊维奇分校) Department of Statistics and Applied Probability, University of California, Santa Barbara, CA, USA(统计学与应用概率系,加州大学圣巴巴拉分校)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏