arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2510.19484 2025-10-23 q-bio.BM cs.AI cs.LG 57%

KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge

Zaifei Yang, Hong Chang, Ruibing Hou, Shiguang Shan, Xilin Chen

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, CAS, China(人工智能安全国家重点实验室,计算技术研究所,中国科学院,中国) University of Chinese Academy of Sciences (CAS), China(中国科学院大学(中国科学院))

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19078 2025-10-23 cs.CV 57%

UniHPR: Unified Human Pose Representation via Singular Value Contrastive Learning

Zhongyu Jiang, Wenhao Chai, Lei Li, Zhuoran Zhou, Cheng-Yen Yang, Jenq-Neng Hwang

机构 * University of Washington(华盛顿大学) University of Copenhagen(哥本哈根大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18053 2025-10-22 cs.LG cs.AI 57%

Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models

Jiajun Fan, Tong Wei, Chaoran Cheng, Yuxin Chen, Ge Liu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18036 2025-10-22 cs.SD cs.LG eess.AS 57%

Transformer Redesign for Late Fusion of Audio-Text Features on Ultra-Low-Power Edge Hardware

Stavros Mitsis, Ermos Hadjikyriakos, Humaid Ibrahim, Savvas Neofytou, Shashwat Raman, James Myles, Eiman Kanjo

机构 * Nottingham Trent University(诺丁汉特伦特大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11741 2025-10-22 cs.AI cs.CR 57%

MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs

Geigh Zollicoffer, Minh Vu, Manish Bhattarai

机构 * Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17847 2025-10-22 cs.CV 57%

CoIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity Optimization

Yichen Yan, Ming Zhong, Qi Zhu, Xiaoling Gu, Jinpeng Chen, Huan Li

机构 * The State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术区(滨江)区块链与数据安全研究院) Hangzhou Dianzi University(杭州电子科技大学) School of Computer Science (National Pilot Software Engineering School), BUPT(计算机学院(国家级试点软件工程学院),北京邮电大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 22 pages, 8 figures, 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17034 2025-10-21 cs.CV 57%

Where, Not What: Compelling Video LLMs to Learn Geometric Causality for 3D-Grounding

Yutong Zhong

机构 * New York University(纽约大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10351 2025-10-21 cs.LG cs.AI 57%

PhysioWave: A Multi-Scale Wavelet-Transformer for Physiological Signal Representation

Yanlong Chen, Mattia Orlandi, Pierangelo Maria Rapa, Simone Benatti, Luca Benini, Yawei Li

机构 * IIS, ETH Zurich(苏黎世联邦理工学院智能信息系统实验室) DEI, University of Bologna(博洛尼亚大学电子工程学院) DIEF, University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学经济工程学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments 43 pages, 17 figures, 17 tables. Accepted by NeurIPS 2025. Code and data are available at: github.com/ForeverBlue816/PhysioWave

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16555 2025-10-21 cs.AI cs.LG 57%

Urban-R1: Reinforced MLLMs Mitigate Geospatial Biases for Urban General Intelligence

Qiongyan Wang, Xingchen Zou, Yutian Jiang, Haomin Wen, Jiaheng Wei, Qingsong Wen, Yuxuan Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Carnegie Mellon University(卡内基梅隆大学) Squirrel Ai Learning

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13316 2025-10-16 cs.CV 57%

Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests

Fitim Abdullahu, Helmut Grabner

机构 * Zurich University of Applied Sciences(苏黎世应用科学大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08163 2025-10-16 cs.CL 57%

ARM2: Adaptive Reasoning Model with Vision Understanding and Executable Code

Jian Xie, Zhendong Chu, Aoxiao Zhong, Kai Zhang, Mingzhe Han, Xing Fan, Jialie Shen, Qingsong Wen

机构 * Squirrel Ai Learning The Ohio State University(俄亥俄州立大学) City St George’s, University of London(伦敦城市圣乔治学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07613 2025-10-16 cs.CV 57%

Data-Efficient Fine-Tuning of Vision-Language Models for Diagnosis of Alzheimer's Disease

Fangqi Cheng, Surajit Ray, Xiaochen Yang

机构 * School of Mathematics and Statistics, University of Glasgow, UK(数学与统计学学院,格拉斯哥大学)

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

Comments Accepted at MICAD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12798 2025-10-15 cs.CV 57%

Detect Anything via Next Point Prediction

Qing Jiang, Junan Huo, Xingyu Chen, Yuda Xiong, Zhaoyang Zeng, Yihao Chen, Tianhe Ren, Junzhi Yu, Lei Zhang

机构 * International Digital Economy Academy (IDEA)(国际数字经济学院(IDEA))

专题命中 多模态训练与对齐 :MLLM(abstract);分类 cs.CV

Comments homepage: https://rex-omni.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11456 2025-10-14 cs.CV 57%

Coupled Degradation Modeling and Fusion: A VLM-Guided Degradation-Coupled Network for Degradation-Aware Infrared and Visible Image Fusion

Tianpei Zhang, Jufeng Zhao, Yiming Zhu, Guangmang Cui

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11449 2025-10-14 cs.CV 57%

Enhancing Maritime Domain Awareness on Inland Waterways: A YOLO-Based Fusion of Satellite and AIS for Vessel Characterization

Geoffery Agorku, Sarah Hernandez, Hayley Hames, Cade Wagner

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09826 2025-10-14 cs.CV 57%

Isolated Channel Vision Transformers: From Single-Channel Pretraining to Multi-Channel Finetuning

Wenyi Lian, Patrick Micke, Joakim Lindblad, Nataša Sladoje

机构 * Department of Information Technology Uppsala University(信息科技系乌普萨拉大学) Department of Immunology, Genetics and Pathology Uppsala University(免疫学、遗传学和病理学系乌普萨拉大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Paper has been accepted by BMVC as an Oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17360 2025-10-14 cs.CV 57%

UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter Tuning

Maoxun Yuan, Bo Cui, Tianyi Zhao, Jiayi Wang, Shan Fu, Xue Yang, Xingxing Wei

机构 * Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) CTTL-Terminal, China Academy of Information and Communications Technology(信息通信技术中国科学院CTTL终端) School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 6 figures, Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23663 2025-10-10 cs.CV 57%

HIVTP: A Training-Free Method to Improve VLMs Efficiency via Hierarchical Visual Token Pruning Using Middle-Layer-Based Importance Score

Jingqi Xu, Jingxi Lu, Chenghao Li, Sreetama Sarkar, Peter A. Beerel

机构 * University of Southern California(美国南加州大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15184 2025-10-08 cs.CV 57%

AuxDet: Auxiliary Metadata Matters for Omni-Domain Infrared Small Target Detection

Yangting Shi, Yinfei Zhu, Renjie He, Le Hui, Meng Cai, Ming-Ming Cheng, Yimian Dai

机构 * School of Electronics and Information(电子信息学院) Northwestern Polytechnical University(西北工业大学) Shaanxi Key Laboratory of Information Acquisition and Processing(陕西省信息获取与处理重点实验室) Luoyang Institute of Electro-Optical Equipment under AVIC(航空工业集团洛阳光电设备研究所) VCIP, College of Computer Science(VCIP,计算机科学学院) Nankai University(南开大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04196 2025-10-07 cs.AI cs.LG 57%

COSMO-RL: Towards Trustworthy LMRMs via Joint Safety and Stability

Yizhuo Ding, Mingkang Chen, Qiuhua Liu, Fenghua Weng, Wanying Qu, Yue Yang, Yugang Jiang, Zuxuan Wu, Yanwei Fu, Wenqi Shao

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学) Shenzhen University(深圳大学) ShanghaiTech University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01454 2025-10-03 cs.CV cs.LG 57%

Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories

Nilay Naharas, Dang Nguyen, Nesihan Bulut, Mohammadhossein Bateni, Vahab Mirrokni, Baharan Mirzasoleiman

机构 * Department of Computer Science, University of California Los Angeles(加州大学洛杉矶分校计算机科学系) Google Research(谷歌研究)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 30 pages, 10 figures, 5 tables, link: https://bigml-cs-ucla.github.io/XMAS-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11997 2025-10-03 cs.LG cs.AI 57%

Can LLMs Find Fraudsters? Multi-level LLM Enhanced Graph Fraud Detection

Tairan Huang, Yili Wang, Qiutong Li, Changlong He, Jianliang Gao

机构 * Central South University(中南大学) Hongkong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19186 2025-10-03 cs.CV cs.RO 57%

LRFusionPR: A Polar BEV-Based LiDAR-Radar Fusion Network for Place Recognition

Zhangshuo Qi, Luqi Cheng, Zijie Zhou, Guangming Xiong

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted by IEEE Robotics and Automation Letters (RAL) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01141 2025-10-02 cs.AI 57%

Apriel-1.5-15b-Thinker

Shruthan Radhakrishna, Aman Tiwari, Aanjaneya Shukla, Masoud Hashemi, Rishabh Maheshwary, Shiva Krishna Reddy Malay, Jash Mehta, Pulkit Pattnaik, Saloni Mittal, Khalil Slimi, Kelechi Ogueji, Akintunde Oladipo, Soham Parikh, Oluwanifemi Bamgbose, Toby Liang, Ahmed Masry, Khyati Mahajan, Sai Rajeswar Mudumba, Vikas Yadav, Sathwik Tejaswi Madhusudhan, Torsten Scholak, Sagar Davasam, Srinivas Sunkara, Nicholas Chapados

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01926 2025-10-02 cs.CV 57%

IC-Custom: Diverse Image Customization via In-Context Learning

Yaowei Li, Xiaoyu Li, Zhaoyang Zhang, Yuxuan Bian, Gan Liu, Xinyuan Li, Jiale Xu, Wenbo Hu, Yating Liu, Lingen Li, Jing Cai, Yuexian Zou, Yancheng He, Ying Shan

机构 * Peking University(北京大学) ARC Lab, Tencent PCG(腾讯PCG ARC实验室) Tencent(腾讯) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Revised version with improved writing. Project page: https://liyaowei-stu.github.io/project/IC_Custom

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01030 2025-10-02 cs.AI 57%

Uncovering the Computational Ingredients of Human-Like Representations in LLMs

Zach Studdiford, Timothy T. Rogers, Kushin Mukherjee, Siddharth Suresh

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Stanford University(斯坦福大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18049 2025-10-02 cs.CV 57%

SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework

Gaole Dai, Menghang Dong, Rongyu Zhang, Ruichuan An, Shanghang Zhang, Tiejun Huang

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26039 2025-10-01 cs.CV 57%

SGS: Segmentation-Guided Scoring for Global Scene Inconsistencies

Gagandeep Singh, Samudi Amarsinghe, Urawee Thani, Ki Fung Wong, Priyanka Singh, Xue Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12737 2025-10-01 cs.CV 57%

PolSAM: Polarimetric Scattering Mechanism Informed Segment Anything Model

Yuqing Wang, Zhongling Huang, Shuxin Yang, Hao Tang, Xiaolan Qiu, Junwei Han, Dingwen Zhang

机构 * the BRain and Artificial INtelligence Lab (BRAIN LAB), School of Automation, Northwestern Polytechnical University(人工智能实验室(BRAIN LAB)、自动化学院、西北工业大学) the Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院) the National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室、计算机科学学院、北京大学) the National Key Laboratory of Microwave Imaging Technology, Chinese Academy of Sciences(微波成像技术国家重点实验室、中国科学院) the Aerospace Information Research Institute, Chinese Academy of Sciences(航空信息研究所、中国科学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments The manuscript is 15 pages long, includes 13 figures and 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24878 2025-09-30 cs.CV cs.RO 57%

ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation

Jiuhong Xiao, Roshan Nayak, Ning Zhang, Daniel Tortei, Giuseppe Loianno

机构 * New York University(纽约大学) Technology Innovation Institute(技术创新研究所) University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 23 pages including the checklist and appendix. Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏