arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2503.12897 2025-07-22 cs.LG cs.AI 57%

Federated Continual Instruction Tuning

Haiyang Guo, Fanhu Zeng, Fei Zhu, Wenzhuo Liu, Da-Han Wang, Jian Xu, Xu-Yao Zhang, Cheng-Lin Liu

机构 * School of Advanced Interdisciplinary Sciences, UCAS(中国科学院大学先进交叉学科学院) MAIS, CASIA(中国科学院自动化所人工智能研究所) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) Centre for Artificial Intelligence and Robotics, HKISI-CAS(香港科技大学人工智能与机器人中心,中国科学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04510 2025-07-22 cs.SE cs.AI 57%

CGP-Tuning: Structure-Aware Soft Prompt Tuning for Code Vulnerability Detection

Ruijun Feng, Hammond Pearce, Pietro Liguori, Yulei Sui

机构 * School of Computer Science and Engineering, University of New South Wales (UNSW)(新南威尔士大学计算机科学与工程学院) Department of Electrical Engineering and Information Technology, University of Naples Federico II(那不勒斯费德里克二世大学电气工程与信息技术系)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

Comments Accepted by IEEE Transactions on Software Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13673 2025-07-21 cs.CV 57%

MaskHOI: Robust 3D Hand-Object Interaction Estimation via Masked Pre-training

Yuechen Xie, Haobo Jiang, Jian Yang, Yigong Zhang, Jin Xie

机构 * PCA Lab, School of Intelligence Science and Technology, Nanjing University(PCA实验室,智能科学与技术学院,南京大学) Nanyang Technological University(南洋理工大学) Nanjing University of Science and Technology(南京理工大学) Nankai University(南开大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 8 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12942 2025-07-18 cs.CV 57%

Weakly Supervised Visible-Infrared Person Re-Identification via Heterogeneous Expert Collaborative Consistency Learning

Yafei Zhang, Lingqi Kong, Huafeng Li, Jie Wen

机构 * Faculty of Information Engineering and Automation, Kunming University of Science and Technology(昆明理工大学信息工程与自动化学院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12807 2025-07-18 cs.CV 57%

Semantic-guided Fine-tuning of Foundation Model for Long-tailed Visual Recognition

Yufei Peng, Yonggang Zhang, Yiu-ming Cheung

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15779 2025-07-18 cs.LG cs.AI 57%

Learning Universal Human Mobility Patterns with a Foundation Model for Cross-domain Data Fusion

Haoxuan Ma, Xishun Liao, Yifan Liu, Qinhua Jiang, Chris Stanford, Shangqing Cao, Jiaqi Ma

机构 * Department of Civil and Environmental Engineering, University of California, Los Angeles(加州大学洛杉矶分校土木与环境工程系) Novateur Research Solutions(Novateur研究解决方案) Department of Civil and Environmental Engineering, University of California, Berkeley(加州大学伯克利分校土木与环境工程系)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12382 2025-07-17 cs.CV 57%

Text-driven Multiplanar Visual Interaction for Semi-supervised Medical Image Segmentation

Kaiwen Huang, Yi Zhou, Huazhu Fu, Yizhe Zhang, Chen Gong, Tao Zhou

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology, China(南京理工大学计算机科学与工程学院) School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院) Institute of High Performance Computing, Agency for Science, Technology and Research, Singapore(新加坡科技研究局高性能计算研究所)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 10 pages; 2 figures; Have been accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23603 2025-07-17 cs.CR cs.AI 57%

SoK: Semantic Privacy in Large Language Models

Baihe Ma, Yanna Jiang, Xu Wang, Guangsheng Yu, Qin Wang, Caijun Sun, Chen Li, Xuelei Qi, Ying He, Wei Ni, Ren Ping Liu

机构 * University of Technology Sydney, Australia(悉尼技术大学) Zhejiang Lab, China(浙江实验室) Northeastern University, China(东北大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15137 2025-07-15 cs.CV 57%

Multispectral Detection Transformer with Infrared-Centric Feature Fusion

Seongmin Hwang, Daeyoung Han, Moongu Jeon

机构 * Artificial Intelligence Graduate School, Gwangju Institute of Science and Technology (GIST)(人工智能研究生院,全州科学技术院(GIST)) School of Electrical Engineering and Computer Science, Gwangju Institute of Science and Technology (GIST)(电气工程与计算机科学学院,全州科学技术院(GIST))

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09209 2025-07-15 cs.CV 57%

Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models

Xiao Liang, Di Wang, Zhicheng Jiao, Ronghan Li, Pengfei Yang, Quan Wang, Tat-Seng Chua

机构 * The Key Laboratory of Smart Human-Computer Interaction and Wearable Technology of Shaanxi Province, Xidian University, China(陕西省智能人机交互与可穿戴技术重点实验室,西安电子科技大学) Warren Alpert Medical School, Brown University, USA(布朗大学沃伦·阿尔珀特医学院) National University of Singapore, Singapore(新加坡国立大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09139 2025-07-15 cs.CV 57%

PoseLLM: Enhancing Language-Guided Human Pose Estimation with MLP Alignment

Dewen Zhang, Tahir Hussain, Wangpeng An, Hayaru Shouno

机构 * Department of Informatics, Graduate School of Informatics and Engineering, The University of Electro-Communications(信息学院,信息工程研究生院,东京电通大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07802 2025-07-14 cs.CV 57%

Synergistic Prompting for Robust Visual Recognition with Missing Modalities

Zhihui Zhang, Luanyuan Dai, Qika Lin, Yunfeng Diao, Guangyin Jin, Yufei Guo, Jing Zhang, Xiaoshuai Hao

机构 * Beijing Institute of Technology(北京理工大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Nanjing University of Science and Technology(南京理工大学) National University of Singapore(新加坡国立大学) Systems Laboratory of Anhui Province, Hefei University of Technology(安徽省系统实验室,合肥工业大学) National Innovative Institute of Defense Technology(国家创新防御技术研究院) Intelligent Science & Technology Academy of CASIC(中国航天科技集团智能科学与技术学院) School of Computer Science, Wuhan University(武汉大学计算机学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00901 2025-07-14 cs.CV 57%

A Decade of Deep Learning for Remote Sensing Spatiotemporal Fusion: Advances, Challenges, and Opportunities

Enzhe Sun, Yongchuan Cui, Peng Liu, Jining Yan

机构 * School of Computer Science, China University of Geosciences (Wuhan)(中国地质大学(武汉)计算机科学学院) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空航天信息研究所) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07527 2025-07-11 cs.CV 57%

MAPEX: Modality-Aware Pruning of Experts for Remote Sensing Foundation Models

Joelle Hanna, Linus Scheibenreif, Damian Borth

机构 * AIML Lab, School of Computer Science University of St.Gallen(人工智能实验室、计算机科学学院、圣加伦大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06203 2025-07-11 cs.CL 57%

A Survey on Latent Reasoning

Rui-Jie Zhu, Tianhao Peng, Tianhao Cheng, Xingwei Qu, Jinfa Huang, Dawei Zhu, Hao Wang, Kaiwen Xue, Xuanliang Zhang, Yong Shan, Tianle Cai, Taylor Kergan, Assel Kembay, Andrew Smith, Chenghua Lin, Binh Nguyen, Yuqi Pan, Yuhong Chou, Zefan Cai, Zhenhe Wu, Yongchi Zhao, Tianyu Liu, Jian Yang, Wangchunshu Zhou, Chujie Zheng, Chongxuan Li, Yuyin Zhou, Zhoujun Li, Zhaoxiang Zhang, Jiaheng Liu, Ge Zhang, Wenhao Huang, Jason Eshraghian

机构 * UCSC(加州大学圣克鲁兹分校) FDU(福建师范大学) NJU(南京大学) PKU(北京大学) RUC(俄罗斯乌拉尔联邦大学) UoM(马来西亚马来亚大学) UW-Madison(威斯康星大学麦迪逊分校) PolyU M-A-P(马大-阿联酋合作项目)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04510 2025-07-08 eess.IV cs.CV 57%

Dynamic Frequency Feature Fusion Network for Multi-Source Remote Sensing Data Classification

Yikang Zhao, Feng Gao, Xuepeng Jin, Junyu Dong, Qian Du

机构 * State Key Laboratory of Physical Oceanography, Ocean University of China(中国海洋大学物理海洋学国家重点实验室) Department of Electrical and Computer Engineering, Mississippi State University(密苏里州立大学电气与计算机工程系)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IEEE GRSL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19281 2025-07-08 eess.SP cs.AI cs.LG 57%

Integrating Biological and Machine Intelligence: Attention Mechanisms in Brain-Computer Interfaces

Jiyuan Wang, Weishan Ye, Jialin He, Li Zhang, Gan Huang, Zhuliang Yu, Zhen Liang

机构 * The School of Biomedical Engineering(生物医学工程学院) The Guangdong Provincial Key Laboratory of Biomedical Measurements(生物医学测量与超声成像广东省重点实验室) Shien-Ming Wu School of Intelligent Engineering(沈铭武智能工程学院) Institute for Super Robotics(超机器人研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03283 2025-07-08 cs.CV 57%

MolVision: Molecular Property Prediction with Vision Language Models

Deepan Adak, Yogesh Singh Rawat, Shruti Vyas

机构 * NIT Kurukshetra(尼特学院) University of Central Florida(中央佛罗里达大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00505 2025-07-08 cs.CV 57%

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs

Haoran Lou, Chunxiao Fan, Ziyan Liu, Yuexin Wu, Xinliang Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beihang University(北航)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00420 2025-07-04 cs.LG cs.CV stat.ML 57%

TAROT: Targeted Data Selection via Optimal Transport

Lan Feng, Fan Nie, Yuejiang Liu, Alexandre Alahi

机构 * EPFL(苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01673 2025-07-03 cs.CV 57%

Facial Emotion Learning with Text-Guided Multiview Fusion via Vision-Language Model for 3D/4D Facial Expression Recognition

Muzammil Behzad

机构 * King Fahd University of Petroleum and Minerals(国王法赫德石油矿物大学) SDAIA-KFUPM Joint Research Center on Artificial Intelligence(SDAIA-KFUPM人工智能联合研究中心)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01368 2025-07-03 cs.CV cs.LG 57%

Activation Reward Models for Few-Shot Model Alignment

Tianning Chai, Chancharik Mitra, Brandon Huang, Gautam Rajendrakumar Gare, Zhiqiu Lin, Assaf Arbelle, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Deva Ramanan, Roei Herzig

机构 * University of California, Berkeley(加州大学伯克利分校) Carnegie Mellon University(卡内基梅隆大学) IBM Research(IBM研究院) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23940 2025-07-02 cs.CL 57%

Graft: Integrating the Domain Knowledge via Efficient Parameter Synergy for MLLMs

Yang Dai, Jianxiang An, Tianwei Lin, Hongyang He, Hongzhe Huang, Wenqiao Zhang, Zheqi Lv, Siliang Tang, Yueting Zhuang

机构 * Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23235 2025-07-01 cs.CL 57%

Generalist Reward Models: Found Inside Large Language Models

Yi-Chen Li, Tian Xu, Yang Yu, Xuqin Zhang, Xiong-Hui Chen, Zhongxiang Ling, Ningjing Chao, Lei Yuan, Zhi-Hua Zhou

机构 * National Key Laboratory for Novel Software Technology, Nanjing University, China(国家新型软件技术重点实验室,南京大学,中国) School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学,中国)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16398 2025-07-01 cs.CV 57%

HyperPath: Knowledge-Guided Hyperbolic Semantic Hierarchy Modeling for WSI Analysis

Peixiang Huang, Yanyan Huang, Weiqin Zhao, Junjun He, Lequan Yu

机构 * The University of Hong Kong(香港大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22624 2025-07-01 cs.CV 57%

Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning

Zuyao You, Zuxuan Wu

机构 * Fudan University(复旦大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22359 2025-06-30 cs.NI cs.AI 57%

Concept-Level AI for Telecom: Moving Beyond Large Language Models

Viswanath Kumarskandpriya, Abdulhalim Dandoush, Abbas Bradai, Ali Belgacem

机构 * Esme Research Lab, SA ESME, Ivry-Sur-Seine, France(Esme研究实验室,SA ESME,法国伊维尔--sur-塞纳) University of Doha for Science and Technology (UDST)(多哈科学技术大学) Côte d'Azur University, LEAT, Sophia Antipolis, France(蔚蓝海岸大学,LEAT,法国索菲亚-安蒂波利斯)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22068 2025-06-30 cs.AI 57%

Query as Test: An Intelligent Driving Test and Data Storage Method for Integrated Cockpit-Vehicle-Road Scenarios

Shengyue Yao, Runqing Guo, Yangyang Qin, Miangbing Meng, Jipeng Cao, Yilun Lin, Yisheng Lv, Fei-Yue Wang

机构 * Peking University International Innovation Center, Lin-gang Special Area (PKU-IICSH)(北京大学国际创新中心,临港特殊区域(PKU-IICSH)) Onesyn (Shanghai) Technology Co., Ltd(上海奥森科技有限公司) CATARC Automotive Technology(Shanghai) Co.,Ltd(CATARC汽车技术(上海)有限公司) ZEEKR Intelligent Technology Holding Limited(ZEKR智能技术控股有限公司) Department of Automation, Tsinghua University(清华大学自动化系) State Key Laboratory for Management and Control of Complex Systems, Chinese Academy of Sciences(复杂系统管理与控制国家重点实验室,中国科学院) Macao Institute of Systems Engineering, Macau University of Science and Technology(澳门系统工程研究院,澳门科技大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Submitted to IEEE Transaction on Vehicular Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21895 2025-06-30 cs.CV 57%

Exploring Task-Solving Paradigm for Generalized Cross-Domain Face Anti-Spoofing via Reinforcement Fine-Tuning

Fangling Jiang, Qi Li, Weining Wang, Gang Wang, Bing Liu, Zhenan Sun

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20986 2025-06-27 cs.CV 57%

EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning

Xiao Zhang, Yongqiang Ma, Haodong Jing, Nanning Zheng

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家级重点实验室) National Engineering Research Center for Visual Information and Applications(视觉信息与应用国家工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所) Xi’an Jiaotong University(西安交通大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏