arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2505.18640 2025-09-30 cs.LG cs.AI 57%

ThanoRA: Task Heterogeneity-Aware Multi-Task Low-Rank Adaptation

Jian Liang, Wenke Huang, Xianda Guo, Guancheng Wan, Bo Du, Mang Ye

机构 * Wuhan University(武汉大学) ByteDance(字节跳动) National University of Defense Technology(国防科技大学) Nanyang Technological University(南洋理工大学) The AGH University of Krakow(克拉科夫AGH大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23772 2025-09-30 cs.CV stat.AP 57%

A Modality-Tailored Graph Modeling Framework for Urban Region Representation via Contrastive Learning

Yaya Zhao, Kaiqi Zhao, Zixuan Tang, Zhiyuan Liu, Xiaoling Lu, Yalei Du

机构 * Center for Applied Statistics, School of Statistics, Innovation Platform, Renmin University of China(应用统计中心、统计学院、创新平台、中国人民大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23641 2025-09-30 cs.CV cs.RO 57%

From Static to Dynamic: a Survey of Topology-Aware Perception in Autonomous Driving

Yixiao Chen, Ruining Yang, Xin Chen, Jia He, Dongliang Xu, Yue Yao

机构 * Sems Shandong University(山东大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23273 2025-09-30 cs.CV 57%

SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction

Yihao Ding, Soyeon Caren Han, Yanbei Jiang, Yan Li, Zechuan Li, Yifan Peng

机构 * The University of Western Australia(西澳大学) The University of Melbourne(墨尔本大学) The University of Sydney(悉尼大学) Weill Cornell Medicine, Cornell University(韦尔·科恩医学中心,康奈尔大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14364 2025-09-29 cs.CL 57%

Position IDs Matter: An Enhanced Position Layout for Efficient Context Compression in Large Language Models

Runsong Zhao, Xin Liu, Xinyu Liu, Pengcheng Huang, Chunyang Xiao, Tong Xiao, Jingbo Zhu

机构 * NLP Lab, School of Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程学院自然语言处理实验室) NiuTrans Research, Shenyang, China(NiuTrans研究院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20240 2025-09-25 cs.LG cs.AI 57%

A HyperGraphMamba-Based Multichannel Adaptive Model for ncRNA Classification

Xin An, Ruijie Li, Qiao Ning, Hui Li, Qian Ma, Shikai Guo

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments 9 pages, 17 figures (including subfigures), 1 table. Xin An and Ruijie Li contributed equally to this work and should be considered co-first authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13739 2025-09-25 cs.CV 57%

Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector

Yiming Cao, Yanjie Li, Kaisheng Liang, Bin Xiao

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19733 2025-09-25 cs.CV 57%

Robust RGB-T Tracking via Learnable Visual Fourier Prompt Fine-tuning and Modality Fusion Prompt Generation

Hongtao Yang, Bineng Zhong, Qihua Liang, Zhiruo Zhu, Yaozong Zheng, Ning Li

机构 * Key Laboratory of Education Blockchain and Intelligent Technology, Ministry of Education, Guangxi Normal University(教育区块链与智能技术重点实验室,教育部,广西师范大学) Guangxi Key Lab of Multi-Source Information Mining and Security, Guangxi Normal University(多源信息挖掘与安全广西重点实验室,广西师范大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by TMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16549 2025-09-25 cs.CV 57%

Efficient Rectified Flow for Image Fusion

Zirui Wang, Jiayi Zhang, Tianwei Guan, Yuhan Zhou, Xingyuan Li, Minjing Dong, Jinyuan Liu

机构 * City University of Hong Kong(香港城市大学) Dalian University of Technology(大连理工大学) Chinese University of Hong Kong(香港中文大学) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19860 2025-09-25 cs.CV cs.LG 57%

SpaRC: Sparse Radar-Camera Fusion for 3D Object Detection

Philipp Wolters, Johannes Gilg, Torben Teepe, Fabian Herzog, Felix Fent, Gerhard Rigoll

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 18 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18743 2025-09-24 cs.CV 57%

TriFusion-AE: Language-Guided Depth and LiDAR Fusion for Robust Point Cloud Processing

Susmit Neogi

机构 * Department of Mechanical Engineering(机械工程系) Indian Institute of Technology Bombay(印度理工学院班加罗尔) Mumbai, India(孟买,印度)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18738 2025-09-24 cs.CV 57%

HyPSAM: Hybrid Prompt-driven Segment Anything Model for RGB-Thermal Salient Object Detection

Ruichao Hou, Xingyuan Li, Tongwei Ren, Dongming Zhou, Gangshan Wu, Jinde Cao

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院) School of Mathematics, Southeast University(东南大学数学学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18733 2025-09-24 cs.CV 57%

Knowledge Transfer from Interaction Learning

Yilin Gao, Kangyi Chen, Zhongxing Peng, Hengjie Lu, Shugong Xu

机构 * Shanghai University(上海大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18613 2025-09-24 cs.CV 57%

MLF-4DRCNet: Multi-Level Fusion with 4D Radar and Camera for 3D Object Detection in Autonomous Driving

Yuzhi Wu, Li Xiao, Jun Liu, Guangfeng Jiang, XiangGen Xia

机构 * MoE Key Laboratory of Brain-Inspired Intelligence Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知教育部重点实验室,中国科学技术大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(人工智能研究院,合肥国家科学中心) Department of Electronic Engineering and Information Science, University of Science and Technology of China(电子工程与信息科学系,中国科学技术大学) Department of Electrical and Computer Engineering, University of Delaware(电气与计算机工程系,德克萨斯大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18152 2025-09-24 cs.LG cs.AI 57%

WLFM: A Well-Logs Foundation Model for Multi-Task and Cross-Well Geological Interpretation

Zhenyu Qi, Qing Yu, Jichen Wang, Yun-Bo Zhao, Zerui Li, Wenjun Lv

机构 * Institute of Advanced Technology, University of Science and Technology of China(科学技术大学先进技术研究所) Department of Automation, University of Science and Technology of China(科学技术大学自动化系) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合国家科学中心人工智能研究所)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19972 2025-09-24 cs.CV 57%

DAOcc: 3D Object Detection Assisted Multi-Sensor Fusion for 3D Occupancy Prediction

Zhen Yang, Yanpeng Dong, Jiayu Wang, Heng Wang, Lichao Ma, Zijian Cui, Qi Liu, Haoran Pei, Kexin Zhang, Chao Zhang

机构 * Beijing Mechanical Equipment Institute(北京机械设备研究所)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments TCSVT Accepted version (not the final published version)

Journal ref IEEE Transactions on Circuits and Systems for Video Technology, 2025, Print ISSN: 1051-8215, Online ISSN: 1558-2205

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17079 2025-09-23 cs.CV 57%

A Dual-Modulation Framework for RGB-T Crowd Counting via Spatially Modulated Attention and Adaptive Fusion

Yuhong Feng, Hongtao Chen, Qi Zhang, Jie Chen, Zhaoxi He, Mingzhe Liu, Jianghai Liao

机构 * College of Computer Science(计算机科学学院) Software Engineering, Shenzhen University, Shenzhen, China(软件工程,深圳大学,深圳,中国)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16348 2025-09-23 cs.AI 57%

A Unified AI Approach for Continuous Monitoring of Human Health and Diseases from Intensive Care Unit to Home with Physiological Foundation Models (UNIPHY+)

Minxiao Wang, Saurabh Kataria, Juntong Ni, Timothy G. Buchman, Jocelyn Grunwell, Mark Mai, Wei Jin, Matthew Clark, Stephanie Brown, Michael Fundora, Puneet Sharma, Tony Pan, Sam Khan, Timothy Ruchti, Naveen Muthu, Kevin Maher, Sivasubramanium V Bhavani, Xiao Hu

机构 * Emory University(埃默里大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15990 2025-09-22 cs.CV 57%

DAFTED: Decoupled Asymmetric Fusion of Tabular and Echocardiographic Data for Cardiac Hypertension Diagnosis

Jérémie Stym-Popper, Nathan Painchaud, Clément Rambour, Pierre-Yves Courand, Nicolas Thome, Olivier Bernard

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 9 pages, Accepted at MIDL 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14739 2025-09-19 cs.CV 57%

FMGS-Avatar: Mesh-Guided 2D Gaussian Splatting with Foundation Model Priors for 3D Monocular Avatar Reconstruction

Jinlong Fan, Bingyu Hu, Xingguang Li, Yuxiang Yang, Jing Zhang

机构 * HangZhou Dianzi University(杭州电子大学) Shenzhen Polytechnic University(深圳职业技术大学) WuHan University(武汉大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14084 2025-09-19 cs.CV 57%

AD-DINOv3: Enhancing DINOv3 for Zero-Shot Anomaly Detection with Anomaly-Aware Calibration

Jingyi Yuan, Jianxiong Ye, Wenkang Chen, Chenqiang Gao

机构 * School of Intelligent Systems Engineering, Sun Yat-Sen University(智能系统工程学院,中山大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14061 2025-09-18 cs.LG cs.AI 57%

Queen Detection in Beehives via Environmental Sensor Fusion for Low-Power Edge Computing

Chiara De Luca, Elisa Donati

机构 * Institute of Neuroinformatics University of Zurich and ETH Zurich(神经信息学研究所(苏黎世大学和苏黎世联邦理工学院)) Digital Society Initiative University of Zurich(数字社会倡议(苏黎世大学))

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18131 2025-09-18 cs.CV 57%

UniPLV: Towards Label-Efficient Open-World 3D Scene Understanding by Regional Visual Language Supervision

Yuru Wang, Pei Liu, Songtao Wang, Zehan Zhang, Xinyan Lu, Changwei Cai, Hao Li, Fu Liu, Peng Jia, Xianpeng Lang

机构 * Li Auto Inc.(力汽车公司) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.01052 2025-09-18 cs.LG cs.AI stat.ML 57%

Joint data imputation and mechanistic modelling for simulating heart-brain interactions in incomplete datasets

Jaume Banus, Maxime Sermesant, Oscar Camara, Marco Lorenzi

机构 * Inria(法国国家信息与自动化技术研究所) Université Côte d’Azur(蔚蓝海岸大学) PhySense(PhySense公司) Department of Information and Communication Technologies(信息与通信技术系) Universitat Pompeu Fabra(庞培法布拉大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21886 2025-09-17 cs.AI cs.LG eess.SP 57%

Efficient Pain Recognition via Respiration Signals: A Single Cross-Attention Transformer Multi-Window Fusion Pipeline

Stefanos Gkikas, Ioannis Kyprakis, Manolis Tsiknakis

机构 * Foundation for Research \& Technology-Hellas Heraklion Greece Foundation for Research \& Technology-Hellas Hellenic Mediterranean University Heraklion Greece Hellenic Mediterranean University

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments arXiv admin note: text overlap with arXiv:2507.21881, arXiv:2507.21875

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21881 2025-09-17 cs.AI 57%

Multi-Representation Diagrams for Pain Recognition: Integrating Various Electrodermal Activity Signals into a Single Image

Stefanos Gkikas, Ioannis Kyprakis, Manolis Tsiknakis

机构 * Foundation for Research \& Technology-Hellas Heraklion Greece Foundation for Research \& Technology-Hellas Hellenic Mediterranean University Heraklion Greece Hellenic Mediterranean University

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments arXiv admin note: text overlap with arXiv:2507.21875

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11817 2025-09-16 cs.CV 57%

MAFS: Masked Autoencoder for Infrared-Visible Image Fusion and Semantic Segmentation

Liying Wang, Xiaoli Zhang, Chuanmin Jia, Siwei Ma

机构 * Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(教育部符号计算与知识工程重点实验室,吉林大学) Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) National Engineering Research Center of Visual Technology, School of Computer Science, Peking University(视觉技术国家工程研究中心,北京大学计算机科学学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Accepted by TIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11476 2025-09-16 cs.CV cs.LG 57%

Modality-Aware Infrared and Visible Image Fusion with Target-Aware Supervision

Tianyao Sun, Dawei Xiang, Tianqi Ding, Xiang Fang, Yijiashun Qi, Zunduo Zhao

机构 * Independent researcher(独立研究者) Dept. of Computer Science Baylor University(计算机科学系 巴里尔大学) Dept. of Computer Science Engineering University of Connecticut(计算机科学工程系 佛罗里达大学) Dept. of Computer Science New York University(计算机科学系 新 york 大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted by 2025 6th International Conference on Computer Vision and Data Mining (ICCVDM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11102 2025-09-16 cs.CV 57%

Filling the Gaps: A Multitask Hybrid Multiscale Generative Framework for Missing Modality in Remote Sensing Semantic Segmentation

Nhi Kieu, Kien Nguyen, Arnold Wiliem, Clinton Fookes, Sridha Sridharan

机构 * School of Electrical Engineering and Robotics, Queensland University of Technology(电气工程与机器人学学院,昆士兰理工大学) Shield AI

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted to DICTA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09085 2025-09-16 cs.CV 57%

IRDFusion: Iterative Relation-Map Difference guided Feature Fusion for Multispectral Object Detection

Jifeng Shen, Haibo Zhan, Xin Zuo, Heng Fan, Xiaohui Yuan, Jun Li, Wankou Yang

机构 * School of Electrical and Information Engineering, Jiangsu University, Zhenjiang, 212013, China(江苏大学电气与信息工程学院) School of Computer Science and Engineering, Jiangsu University of Science and Technology, Zhenjiang, 212003, China(江苏科技大学计算机科学与工程学院) School of Automation, Southeast University, Nanjing, 210096, China(东南大学自动化学院) Department of Computer Science and Engineering, University of North Texas, Denton, TX 76207, USA(北卡罗来纳州立大学计算机科学与工程系) School of Computing, Nanjing Normal University, Nanjing, 210046, Jiangsu(南京师范大学计算机学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 31 pages,6 figures, submitted on 3 Sep,2025

详情

展开后加载摘要…

URL PDF HTML 收藏