arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2506.02433 2025-06-04 cs.CV 57%

Empowering Functional Neuroimaging: A Pre-trained Generative Framework for Unified Representation of Neural Signals

Weiheng Yao, Xuhang Chen, Shuqiang Wang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12599 2025-06-04 cs.AI cs.LG 57%

Kimi k1.5: Scaling Reinforcement Learning with LLMs

Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, Chuning Tang, Congcong Wang, Dehao Zhang, Enming Yuan, Enzhe Lu, Fengxiang Tang, Flood Sung, Guangda Wei, Guokun Lai, Haiqing Guo, Han Zhu, Hao Ding, Hao Hu, Hao Yang, Hao Zhang, Haotian Yao, Haotian Zhao, Haoyu Lu, Haoze Li, Haozhen Yu, Hongcheng Gao, Huabin Zheng, Huan Yuan, Jia Chen, Jianhang Guo, Jianlin Su, Jianzhou Wang, Jie Zhao, Jin Zhang, Jingyuan Liu, Junjie Yan, Junyan Wu, Lidong Shi, Ling Ye, Longhui Yu, Mengnan Dong, Neo Zhang, Ningchen Ma, Qiwei Pan, Qucheng Gong, Shaowei Liu, Shengling Ma, Shupeng Wei, Sihan Cao, Siying Huang, Tao Jiang, Weihao Gao, Weimin Xiong, Weiran He, Weixiao Huang, Weixin Xu, Wenhao Wu, Wenyang He, Xianghui Wei, Xianqing Jia, Xingzhe Wu, Xinran Xu, Xinxing Zu, Xinyu Zhou, Xuehai Pan, Y. Charles, Yang Li, Yangyang Hu, Yangyang Liu, Yanru Chen, Yejie Wang, Yibo Liu, Yidao Qin, Yifeng Liu, Ying Yang, Yiping Bao, Yulun Du, Yuxin Wu, Yuzhi Wang, Zaida Zhou, Zhaoji Wang, Zhaowei Li, Zhen Zhu, Zheng Zhang, Zhexu Wang, Zhilin Yang, Zhiqi Huang, Zihao Huang, Ziyao Xu, Zonghan Yang, Zongyu Lin

机构 * Kimi Team(Kimi 团队)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01203 2025-06-03 cs.CV 57%

Self-Supervised Multi-View Representation Learning using Vision-Language Model for 3D/4D Facial Expression Recognition

Muzammil Behzad

机构 * King Fahd University of Petroleum and Minerals(国王法赫德石油矿物大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03801 2025-06-03 cs.LG cs.AI 57%

Semantic-guided Representation Learning for Multi-Label Recognition

Ruhui Zhang, Hezhe Qiao, Pengcheng Xu, Mingsheng Shang, Lin Chen

机构 * Chongqing University of Posts and Telecommuncation, China(重庆邮电大学) Chongqing Institute of Green and Intelligent Technology, Chinese Academy of Sciences, China(重庆绿色智能技术研究院,中国科学院) Singapore Management University, Singapore(新加坡国立大学)

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.AI

Comments Accepted in ICME2025 Oral (15% of all submissions)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14106 2025-06-02 cs.AI 57%

Advancing Molecular Graph-Text Pre-training via Fine-grained Alignment

Yibo Li, Yuan Fang, Mengmei Zhang, Chuan Shi

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Singapore Management University(新加坡国立大学) China Telecom Bestpay(中国电信最佳支付)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments Accepted by KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21494 2025-05-28 cs.CV 57%

Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment

Xiaojun Jia, Sensen Gao, Simeng Qin, Tianyu Pang, Chao Du, Yihao Huang, Xinfeng Li, Yiming Li, Bo Li, Yang Liu

机构 * Nanyang Technological University, Singapore(南洋理工大学,新加坡) MBZUAI, United Arab Emirates(马克斯·普朗克人工智能研究所,阿拉伯联合酋长国) Sea AI Lab, Singapore(海思人工智能实验室,新加坡) University of Illinois Urbana-Champaign, USA(伊利诺伊大学厄巴纳-香槟分校,美国)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00429 2025-05-28 cs.CV 57%

DADM: Dual Alignment of Domain and Modality for Face Anti-spoofing

Jingyi Yang, Xun Lin, Zitong Yu, Liepiao Zhang, Xin Liu, Hui Li, Xiaochen Yuan, Xiaochun Cao

机构 * Great Bay University(大西洋大学) University of Science and Technology of China(中国科学技术大学) Dongguan Key Laboratory for Intelligence and Information Technology(东莞智能与信息技术重点实验室) GRGBanking Equipment Co., Ltd.(GRGBanking设备有限公司) South China University of Technology(华南理工大学) Lappeenranta University of Technology(拉普兰塔理工大学) Macao Polytechnic University(澳门理工学院) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 18 pages, 9 figures, Code: https://github.com/yjyddq/DADM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20777 2025-05-28 cs.CV 57%

TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs

Zhehan Kan, Yanlin Liu, Kun Yin, Xinghua Jiang, Xin Li, Haoyu Cao, Yinsong Liu, Deqiang Jiang, Xing Sun, Qingmin Liao, Wenming Yang

机构 * Tencent YouTu Lab(腾讯YouTu实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00396 2025-05-28 cs.CV 57%

SPF-Portrait: Towards Pure Text-to-Portrait Customization with Semantic Pollution-Free Fine-Tuning

Xiaole Xian, Zhichao Liao, Qingyu Li, Wenyu Qin, Pengfei Wan, Weicheng Xie, Long Zeng, Linlin Shen, Pingfa Feng

机构 * Shenzhen University(深圳大学) Tsinghua University(清华大学) Kuaishou Technology(快手科技)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05721 2025-05-27 cs.CV 57%

Semantic-Space-Intervened Diffusive Alignment for Visual Classification

Zixuan Li, Lei Meng, Guoqing Chao, Wei Wu, Xiaoshuo Yan, Yimeng Yang, Zhuang Qi, Xiangxu Meng

机构 * School of Software, Shandong University(软件学院,山东大学) School of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术学院,哈尔滨工业大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13617 2025-05-27 cs.CV 57%

Compile Scene Graphs with Reinforcement Learning

Zuyao Chen, Jinlin Wu, Zhen Lei, Marc Pollefeys, Chang Wen Chen

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.08190 2025-05-27 cs.CV cs.LG 57%

Graph Neural Networks for Knowledge Enhanced Visual Representation of Paintings

Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Marcel Worring, Nachoem Wijnberg

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Published in the 29th ACM International Conference on Multimedia (MM '21). This is the camera-ready version. 10 pages, 4 figures

Journal ref Proc. 29th ACM Int. Conf. on Multimedia (MM '21), 2021, pp. 3710-3719

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17994 2025-05-26 cs.CV 57%

Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation

Zhihua Liu, Amrutha Saseendran, Lei Tong, Xilin He, Fariba Yousefi, Nikolay Burlutskiy, Dino Oglic, Tom Diethe, Philip Teare, Huiyu Zhou, Chen Jin

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17821 2025-05-26 cs.CV 57%

ICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification

Shihao Li, Chenglong Li, Aihua Zheng, Jin Tang, Bin Luo

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Multimedia (TMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17110 2025-05-26 cs.CL 57%

Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling

Junlin Li, Guodong DU, Jing Li, Sim Kuan Goh, Wenya Wang, Yequan Wang, Fangming Liu, Ho-Kin Tang, Saleh Alharbi, Daojing He, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) Xiamen University Malaysia(厦门大学马来西亚分校) Nanyang Technological University(南洋理工大学) Beijing Academy of Artificial Intelligence, China(北京人工智能研究院) Peng Cheng Laboratory, China(鹏城实验室) Shaqra University, Saudi Arabia(沙迦大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17040 2025-05-26 cs.LG cs.CL 57%

Generalizing Large Language Model Usability Across Resource-Constrained

Yun-Da Tsai

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments Doctoral disstertation

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14043 2025-05-26 cs.CV 57%

Selective Structured State Space for Multispectral-fused Small Target Detection

Qianqian Zhang, WeiJun Wang, Yunxing Liu, Li Zhou, Hao Zhao, Junshe An, Zihan Wang

机构 * Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学) National Space Science Center, Chinese Academy of Sciences(中国科学院国家空间科学中心) School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院) School of Astronomy and Space Science, University of Chinese Academy of Sciences(中国科学院大学天文与空间科学学院) School of Computing, National University of Singapore(新加坡国立大学计算机学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments This work was submitted to CVPR 2025, but was rejected after being reviewed by 7 reviewers. After revision, it is currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15875 2025-05-23 cs.CV cs.LG 57%

Decouple and Orthogonalize: A Data-Free Framework for LoRA Merging

Shenghe Zheng, Hongzhi Wang, Chenyu Huang, Xiaohui Wang, Tao Chen, Jiayuan Fan, Shuyue Hu, Peng Ye

机构 * Harbin Institute of Technology(哈尔滨工业大学) Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15222 2025-05-23 cs.CV 57%

Continuous Representation Methods, Theories, and Applications: An Overview and Perspectives

Yisi Luo, Xile Zhao, Deyu Meng

机构 * School of Mathematics and Statistics, Xi’an Jiaotong University, Xi’an 710000, China(西安交通大学数学与统计学院) School of Mathematical Sciences, University of Electronic Science and Technology of China, Chengdu 610000, China(电子科技大学数学科学学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06846 2025-05-23 cs.LG cs.AI q-bio.BM 57%

Prot2Chat: Protein LLM with Early-Fusion of Text, Sequence and Structure

Zhicong Wang, Zicheng Ma, Ziqiang Cao, Changlong Zhou, Jun Zhang, Yiqin Gao

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15816 2025-05-22 cs.CV 57%

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM

Penghao Wu, Lewei Lu, Ziwei Liu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15491 2025-05-22 cs.CV 57%

Spectral-Aware Global Fusion for RGB-Thermal Semantic Segmentation

Ce Zhang, Zifu Wan, Simon Stepputtis, Katia Sycara, Yaqi Xie

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15401 2025-05-22 cs.CV 57%

Visual Question Answering on Multiple Remote Sensing Image Modalities

Hichem Boussaid, Lucrezia Tosato, Flora Weissgerber, Camille Kurtz, Laurent Wendling, Sylvain Lobry

机构 * LIPADE, Université Paris Cité, France(巴黎cité大学LIPADE研究所) ONERA, France(法国ONERA研究所)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments EARTHVISION 2025 8 pages, 1 page of supplementary material, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15116 2025-05-22 cs.LG cs.AI cs.SI 57%

Graph Foundation Models: A Comprehensive Survey

Zehong Wang, Zheyuan Liu, Tianyi Ma, Jiazheng Li, Zheyuan Zhang, Xingbo Fu, Yiyang Li, Zhengqing Yuan, Wei Song, Yijun Ma, Qingkai Zeng, Xiusi Chen, Jianan Zhao, Jundong Li, Meng Jiang, Pietro Lio, Nitesh Chawla, Chuxu Zhang, Yanfang Ye

机构 * University of Notre Dame(诺丁汉大学) University of Connecticut(康涅狄格大学) University of Virginia(弗吉尼亚大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Cambridge(剑桥大学) Mila - Québec AI Institute(魁北克AI研究院) Université de Montréal(蒙特利尔大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Github Repo: https://github.com/Zehong-Wang/Awesome-Foundation-Models-on-Graphs. 93 pages, 438 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06897 2025-05-21 cs.CV 57%

Multi-granular body modeling with Redundancy-Free Spatiotemporal Fusion for Text-Driven Motion Generation

Xingzu Zhan, Chen Xie, Honghang Chen, Haoran Sun, Xiaochun Mai

机构 * Shenzhen University(深圳大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 15pages,5figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13081 2025-05-20 cs.LG cs.CV 57%

Walking the Tightrope: Disentangling Beneficial and Detrimental Drifts in Non-Stationary Custom-Tuning

Xiaoyu Yang, Jie Lu, En Yu

机构 * Australian Artificial Intelligence Institute (AAII)(澳大利亚人工智能研究所) Faulty of Engineering and Information Technology(工程与信息技术学院) University of Technology Sydney(悉尼技术大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 17 pages, 5figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11878 2025-05-20 cs.LG cs.AI q-bio.MN 57%

AdaptMol: Adaptive Fusion from Sequence String to Topological Structure for Few-shot Drug Discovery

Yifan Dai, Xuanbai Ren, Tengfei Ma, Qipeng Yan, Yiping Liu, Yuansheng Liu, Xiangxiang Zeng

机构 * College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院) School of Biomedical Science, Hunan University(湖南大学生物医学科学学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17207 2025-05-19 cs.CV 57%

Self-Supervised Representation Learning for Nerve Fiber Distribution Patterns in 3D-PLI

Alexander Oberstrass, Sascha E. A. Muenzing, Meiqi Niu, Nicola Palomero-Gallagher, Christian Schiffer, Markus Axer, Katrin Amunts, Timo Dickscheid

机构 * Institute of Neuroscience and Medicine (INM-1), Research Centre Jülich, Germany(德国汝拉赫研究中心神经科学与医学研究所) Helmholtz AI, Research Centre Jülich, Germany(德国汝拉赫研究中心海德堡人工智能研究所) Cécile & Oskar Vogt Institute of Brain Research, Medical Faculty and University Hospital Düsseldorf, Heinrich Heine University Düsseldorf, Germany(德国杜塞尔多夫海因里希-海涅大学杜塞尔多夫医学院和大学医院塞西尔与奥斯特尔研究所) Department of Physics, University of Wuppertal, Germany(德国乌尔姆大学物理系) Institute of Computer Science, Faculty of Mathematics and Natural Sciences, Heinrich Heine University Düsseldorf, Germany(德国杜塞尔多夫海因里希-海涅大学数学与自然科学学院计算机科学研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Journal version

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06604 2025-05-19 cs.CL 57%

Do we really have to filter out random noise in pre-training data for language models?

Jinghan Ru, Yuxin Xie, Xianwei Zhuang, Yuguo Yin, Zhihui Guo, Zhiming Liu, Qianli Ren, Yuexian Zou

机构 * School of Electronic and Computer Engineering, Peking University(北京理工大学电子与计算机工程学院) University of Electronic Science and Technology of China(电子科技大学) Hong Kong University of Science and Technology(香港理工大学) Sichuan University(四川大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00646 2025-05-16 cs.CL 57%

Phase Diagram of Vision Large Language Models Inference: A Perspective from Interaction across Image and Instruction

Houjing Wei, Yuting Shi, Naoya Inoue

机构 * Japan Advanced Institute of Science and Technology(日本先进科学研究院) RIKEN(日本资源技术研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments 6 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏