arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-12 至 2025-08-12 共收录 100 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 24 篇

2502.20988 2025-08-12 cs.AI cs.CL 62%

Reviewing Clinical Knowledge in Medical Large Language Models: Training and Beyond

Qiyuan Li, Haijiang Liu, Caicai Guo, Chao Gao, Deyu Chen, Meng Wang, Feng Gao, Frank van Harmelen, Jinguang Gu

机构 * School of Computer Science and Technology, Wuhan University of Science and Technology(计算机科学与技术学院,武汉科技大学) Hubei Province Key Laboratory of Intelligent Information Processing and Real-time Industrial System(湖北省智能信息处理与实时工业系统重点实验室) School of Computer Science and Technology, Huazhong University of Science and Technology(计算机科学与技术学院,华中科技大学) School of Cyber Science and Engineering, Wuhan University(网络科学与工程学院,武汉大学) Department of Computer Science, Vrije Universiteit Amsterdam(计算机科学系,阿姆斯特丹自由大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted for publication in Knowledge-Based Systems. The arXiv version is the pre-peer-review preprint, and the final published version is not available here due to publisher policy

Journal ref Knowledge-Based Systems, 114215(2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08252 2025-08-12 cs.CV 57%

ReferSplat: Referring Segmentation in 3D Gaussian Splatting

Shuting He, Guangquan Jie, Changshuo Wang, Yun Zhou, Shuming Hu, Guanbin Li, Henghui Ding

机构 * MoE Key Laboratory of Interdisciplinary Research of Computation and Economics(跨学科计算与经济学研究实验室) Shanghai University of Finance(上海财经大学) Institute of Big Data, College of Computer Science and Artificial Intelligence(大数据研究所,计算机科学与人工智能学院) Fudan University, Shanghai, China(复旦大学,上海,中国) Nanyang Technological University, Singapore(南洋理工大学,新加坡) Sun Yat-sen University, Guangzhou, China(中山大学,广州,中国)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments ICML 2025 Oral, Code: https://github.com/heshuting555/ReferSplat

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07838 2025-08-12 cs.CV 57%

CBDES MoE: Hierarchically Decoupled Mixture-of-Experts for Functional Modules in Autonomous Driving

Qi Xiang, Kunsong Shi, Zhigui Lin, Lei He

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07812 2025-08-12 cs.CV 57%

Semi-supervised Multiscale Matching for SAR-Optical Image

Jingze Gai, Changchun Li

机构 * Nanyang Technological University(南洋理工大学) Jilin University(吉林大学)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV

Comments 15 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05016 2025-08-12 cs.CV eess.IV 57%

AU-IQA: A Benchmark Dataset for Perceptual Quality Assessment of AI-Enhanced User-Generated Content

Shushi Wang, Chunyi Li, Zicheng Zhang, Han Zhou, Wei Dong, Jun Chen, Guangtao Zhai, Xiaohong Liu

机构 * Shanghai Jiao Tong University(上海交通大学) McMaster University(麦斯特大学) Suzhou Key Laboratory of Artificial Intelligence(苏州人工智能重点实验室)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACMMM 2025 Datasets Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06624 2025-08-12 cs.CV 57%

VL-MedGuide: A Visual-Linguistic Large Model for Intelligent and Explainable Skin Disease Auxiliary Diagnosis

Kexin Yu, Zihan Xu, Jialei Xie, Carter Adams

机构 * Jiangsu Ocean University(江苏海洋大学) Federal University of Bahia(巴伊亚联邦大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01416 2025-08-12 cs.RO cs.CV 57%

UniCalib: Targetless LiDAR-Camera Calibration via Probabilistic Flow on Unified Depth Representations

Shu Han, Xubo Zhu, Ji Wu, Ximeng Cai, Wen Yang, Huai Yu, Gui-Song Xia

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) School of Electronic Information, Wuhan University(武汉大学电子信息学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV

Comments 8 pages,5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04305 2025-08-12 eess.IV 50%

Edge2Prompt: Modality-Agnostic Model for Out-of-Distribution Liver Segmentation

Nathan Hollet, Oumeymah Cherkaoui, Philippe C. Cattin, Sidaty El Hadramy

专题命中 多模态评测 :multi-modal(abstract)

Comments 8 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06841 2025-08-12 cs.NE 50%

Memory Enhanced Fractional-Order Dung Beetle Optimization for Photovoltaic Parameter Identification

Yiwei Li, Zhihua Allen-Zhao, Yuncheng Xu, Sanyang Liu

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 7 篇

2403.09333 2025-08-12 cs.CV cs.AI 81%

Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Yufei Zhan, Shurong Zheng, Yousong Zhu, Hongyin Zhao, Fan Yang, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室) Wuhan AI Research, Wuhan, China(武汉人工智能研究所)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025. Codes and datasets are released at https://github.com/jefferyZhan/Griffon

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08137 2025-08-12 cs.LG cs.AI cs.SY eess.SY 79%

MuaLLM: A Multimodal Large Language Model Agent for Circuit Design Assistance with Hybrid Contextual Retrieval-Augmented Generation

Pravallika Abbineni, Saoud Aldowaish, Colin Liechty, Soroosh Noorzad, Ali Ghazizadeh, Morteza Fayazi

机构 * University of Utah(犹他大学) University of Michigan(密歇根大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07621 2025-08-12 cs.CV cs.AI 62%

SOFA: Deep Learning Framework for Simulating and Optimizing Atrial Fibrillation Ablation

Yunsung Chung, Chanho Lim, Ghassan Bidaoui, Christian Massad, Nassir Marrouche, Jihun Hamm

机构 * Department of Computer Science, Tulane University(计算机科学系, Tulane大学) School of Medicine, Tulane University(医学院, Tulane大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at MICCAI 2025. This is the author's original preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07466 2025-08-12 cs.AI 57%

Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs

Dom Huh, Prasant Mohapatra

机构 * UC Davis(加州大学戴维斯分校) University of South Florida(佛罗里达州立大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07010 2025-08-12 cs.MM cs.HC cs.MA 57%

Narrative Memory in Machines: Multi-Agent Arc Extraction in Serialized TV

Roberto Balestri, Guglielmo Pescatore

专题命中 多模态Agent :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07141 2025-08-12 cs.HC 50%

SketchConcept: Sketching-based Concept Recomposition for Product Design using Generative AI

Runlin Duan, Chenfei Zhu, Yuzhao Chen, Dizhi Ma, Jingyu Shi, Ziyi Liu, Karthik Ramani

专题命中 多模态Agent :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05554 2025-08-12 cs.RO 50%

MultiNash-PF: A Particle Filtering Approach for Computing Multiple Local Generalized Nash Equilibria in Trajectory Games

Maulik Bhatt, Iman Askari, Yue Yu, Ufuk Topcu, Huazhen Fang, Negar Mehr

机构 * University of California, Berkeley(加州大学伯克利分校) University of Kansas(堪萨斯大学) University of Minnesota(明尼苏达大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 多模态Agent :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 16 篇

2508.07803 2025-08-12 cs.CV 83%

MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks

Yushen Xu, Xiaosong Li, Zhenyu Kuang, Xiaoqi Cheng, Haishu Tan, Huafeng Li

机构 * School of Physics and Optoelectronic Engineering(物理与光电工程学院) Guangdong-HongKong-Macao Joint Laboratory for Intelligent Micro-Nano Optoelectronic Technology(粤港澳联合智能微纳光电技术实验室) School of Information Engineering and Automation(信息工程与自动化学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07681 2025-08-12 cs.LG cs.AI 83%

MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation

Yooseok Lim, ByoungJun Jeon, Seong-A Park, Jisoo Lee, Sae Won Choi, Chang Wook Jeong, Ho-Geol Ryu, Hongyeol Lee, Hyun-Lim Yang

机构 * Seoul National University Hospital(首尔国立大学医院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06701 2025-08-12 cs.CV cs.AI cs.CL cs.LG cs.SD eess.AS 83%

MMFformer: Multimodal Fusion Transformer Network for Depression Detection

Md Rezwanul Haque, Md. Milon Islam, S M Taslim Uddin Raju, Hamdi Altaheri, Lobna Nassar, Fakhri Karray

机构 * Centre for Pattern Analysis and Machine Intelligence, Department of Electrical and Computer Engineering, University of Waterloo(模式分析与机器智能中心,电气与计算机工程系,滑铁卢大学) School of Engineering and Computing, Department of Computer Science and Engineering, American University of Ras Al Khaimah(工程与计算学院,计算机科学与工程系,阿联酋拉线哈姆斯美国大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted for the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00425 2025-08-12 cs.CV cs.AI 82%

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization

JiangYong Yu, Sifan Zhou, Dawei Yang, Shuo Wang, Shuoyu Li, Xing Hu, Chen Xu, Zukang Xu, Changyong Shu, Zhihang Yuan

机构 * Southeast University(东南大学) Xi'an Jiaotong University(西安交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM MM 2025. First PTQ solution for Multimodal large language models applicable to 5 mainstream MLLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06895 2025-08-12 cs.CV cs.AI 81%

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models

Jianting Tang, Yubo Wang, Haoyu Cao, Linli Xu

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00826 2025-08-12 cs.CL cs.AI cs.LG 81%

HERGC: Heterogeneous Experts Representation and Generative Completion for Multimodal Knowledge Graphs

Yongkang Xiao, Rui Zhang

机构 * University of Minnesota(明尼苏达大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04770 2025-08-12 cs.LG cs.AI q-bio.MN 79%

Bidirectional Hierarchical Protein Multi-Modal Representation Learning

Xuefeng Liu, Songhao Jiang, Chih-chan Tien, Jinbo Xu, Rick Stevens

机构 * Argonne National Laboratory(阿贡国家实验室)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06496 2025-08-12 cs.CV cs.MA 79%

Med-GRIM: Enhanced Zero-Shot Medical VQA using prompt-embedded Multimodal Graph RAG

Rakesh Raj Madavan, Akshat Kaimal, Hashim Faisal, Chandrakala S

机构 * Shiv Nadar University Chennai(施瓦斯纳大学钦奈)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07951 2025-08-12 cs.CV 79%

Scaling Laws for Native Multimodal Models

Mustafa Shukor, Enrico Fini, Victor Guilherme Turrisi da Costa, Matthieu Cord, Joshua Susskind, Alaaeldin El-Nouby

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2025 (Oral). 28 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07536 2025-08-12 cs.LG 78%

Physics-Informed Multimodal Bearing Fault Classification under Variable Operating Conditions using Transfer Learning

Tasfiq E. Alam, Md Manjurul Ahsan, Shivakumar Raman

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06566 2025-08-12 cs.CV cs.AI 73%

Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features

Manish Kansana, Elias Hossain, Shahram Rahimi, Noorbakhsh Amiri Golilarz

机构 * Department of Computer Science and Engineering, Mississippi State University(计算机科学与工程系,密苏里州立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06826 2025-08-12 cs.HC 67%

AdjustAR: AI-Driven In-Situ Adjustment of Site-Specific Augmented Reality Content

Nels Numan, Jessica Van Brummelen, Ziwen Lu, Anthony Steed

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract)

Comments 4 pages, 1 figure, ACM UIST 2025 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07804 2025-08-12 cs.CV 57%

Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning

Bao Li, Xiaomei Zhang, Miao Xu, Zhaoxin Fan, Xiangyu Zhu, Zhen Lei

机构 * CASIA(中国科学院自动化研究所) UCAS(中国科学院大学) CAIR, HKISI, CAS(中国科学院自动化研究所) Beihang University(北京航空航天大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07480 2025-08-12 eess.SP cs.AI cs.LG 57%

EEG-Language Pretraining for Highly Label-Efficient Clinical Phenotyping

Sam Gijsen, Kerstin Ritter

机构 * Charité – Universitätsmedizin Berlin, Department of Psychiatry and Psychotherapy, Berlin, Germany(柏林查理医院医学大学精神病与心理治疗系) Hertie Institute for AI in Brain Health, University of Tübingen, Germany(图宾根大学健康人工智能研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏