arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-04 至 2025-08-04 共收录 45 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2508.00171 2025-08-04 cs.CV cs.CL 82%

On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI

David Restrepo, Ira Ktena, Maria Vakalopoulou, Stergios Christodoulidis, Enzo Ferrante

机构 * MICS, CentraleSupélec - Université Paris-Saclay, France(MICS,中央圣艾尔布兰大学-巴黎萨克雷大学,法国) Google DeepMind, London, UK(谷歌DeepMind,伦敦,英国) CONICET, Universidad de Buenos Aires, Argentina(CONICET,布宜诺斯艾利斯大学,阿根廷)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to MICCAI 2025 1st Workshop on Multimodal Large Language Models (MLLMs) in Clinical Practice

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22062 2025-08-04 cs.CV cs.CL 73%

Meta CLIP 2: A Worldwide Scaling Recipe

Yung-Sung Chuang, Yang Li, Dong Wang, Ching-Feng Yeh, Kehan Lyu, Ramya Raghavendra, James Glass, Lifei Huang, Jason Weston, Luke Zettlemoyer, Xinlei Chen, Zhuang Liu, Saining Xie, Wen-tau Yih, Shang-Wen Li, Hu Xu

机构 * FAIR, Meta(FAIR与Meta公司) MIT(麻省理工学院) Princeton University(普林斯顿大学) New York University(纽约大学)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00043 2025-08-04 cs.CV cs.AI 73%

MR-CLIP: Efficient Metadata-Guided Learning of MRI Contrast Representations

Mehmet Yigit Avci, Pedro Borges, Paul Wright, Mehmet Yigitsoy, Sebastien Ourselin, Jorge Cardoso

机构 * School of Biomedical Engineering and Imaging Sciences, King’s College London, London, UK(生物医学工程与成像科学学院,伦敦国王学院,伦敦,英国) deepc GMBH, Munich, Germany(deepc 德国慕尼黑分公司)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00356 2025-08-04 cs.CV cs.MA 57%

Analyze-Prompt-Reason: A Collaborative Agent-Based Framework for Multi-Image Vision-Language Reasoning

Angelos Vlachos, Giorgos Filandrianos, Maria Lymperaiou, Nikolaos Spanos, Ilias Mitsouras, Vasileios Karampinis, Athanasios Voulodimos

机构 * Artificial Intelligence and Learning Systems Laboratory, National Technical University of Athens(人工智能与学习系统实验室,国家技术大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21053 2025-08-04 cs.LG cs.RO 50%

Flow Matching Policy Gradients

David McAllister, Songwei Ge, Brent Yi, Chung Min Kim, Ethan Weber, Hongsuk Choi, Haiwen Feng, Angjoo Kanazawa

机构 * UC Berkeley(伯克利大学) Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)

专题命中 图文多模态 :multimodal(abstract)

Comments See our blog post at https://flowreinforce.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23859 2025-08-04 astro-ph.IM 50%

radio-llava: Advancing Vision-Language Models for Radio Astronomical Source Analysis

S. Riggi, T. Cecconello, A. Pilzer, S. Palazzo, N. Gupta, A. M. Hopkins, C. Trigilio, G. Umana

专题命中 图文多模态 :multimodal(abstract)

Comments 19 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 6 篇

2508.00632 2025-08-04 cs.AI cs.MA cs.MM 84%

Multi-Agent Game Generation and Evaluation via Audio-Visual Recordings

Alexia Jolicoeur-Martineau

机构 * Samsung SAIL Montréal(三星SAIL蒙特利尔)

专题命中 音频语音多模态 :audio-visual(title,abstract);omni-modal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00760 2025-08-04 cs.CL cs.AI 81%

MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection under Cloaking Perturbations

Qiyao Xue, Yuchen Dou, Ryan Shi, Xiang Lorraine Li, Wei Gao

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00784 2025-08-04 cs.AI 79%

Unraveling Hidden Representations: A Multi-Modal Layer Analysis for Better Synthetic Content Forensics

Tom Or, Omri Azencot

机构 * Ben Gurion University of the Negev(本· Gurion 内盖夫大学)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00391 2025-08-04 cs.CV eess.AS 62%

Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition

Guanjie Huang, Danny H. K. Tsang, Shan Yang, Guangzhi Lei, Li Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tencent AI Lab(腾讯AI实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、eess.AS

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00160 2025-08-04 cs.HC cs.AI cs.SD eess.AS 62%

DeformTune: A Deformable XAI Music Prototype for Non-Musicians

Ziqing Xu, Nick Bryan-Kinns

机构 * Creative Computing Institute, University of the Arts London(创意计算研究所,伦敦艺术大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments In Proceedings of Explainable AI for the Arts Workshop 2025 (XAIxArts 2025) arXiv:2406.14485

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00205 2025-08-04 cs.CV 57%

Learning Personalised Human Internal Cognition from External Expressive Behaviours for Real Personality Recognition

Xiangyu Kong, Hengde Zhu, Haoqin Sun, Zhihao Guo, Jiayan Gu, Xinyi Ni, Wei Zhang, Shizhe Liu, Siyang Song

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 4 篇

2507.06603 2025-08-04 cs.CV 79%

Cross-Modal Dual-Causal Learning for Long-Term Action Recognition

Xu Shaowu, Jia Xibin, Gao Junyu, Sun Qianmei, Chang Jing, Fan Chao

机构 * Beijing University of Technology(北京理工大学) Chinese Academy of Sciences(中国科学院) Capital Medical University(首都医科大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00210 2025-08-04 stat.CO stat.ME 78%

Efficient rare event estimation for multimodal and high-dimensional system reliability via subset adaptive importance sampling

Sara Helal, Victor Elvira

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00085 2025-08-04 cs.CV cs.AI 62%

Punching Bag vs. Punching Person: Motion Transferability in Videos

Raiyaan Abdullah, Jared Claypoole, Michael Cogswell, Ajay Divakaran, Yogesh Rawat

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) Center for Vision Technology, SRI International(视觉技术中心,SRI国际)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02713 2025-08-04 cs.CV cs.CL 62%

LLaVA-Video: Video Instruction Tuning With Synthetic Data

Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, Chunyuan Li

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) BUPT(北京邮电大学) ByteDance(字节跳动)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Project page: https://llava-vl.github.io/blog/2024-09-30-llava-video/; Accepted at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 3 篇

2508.00332 2025-08-04 cs.CL 79%

Improving Multimodal Contrastive Learning of Sentence Embeddings with Object-Phrase Alignment

Kaiyan Zhao, Zhongtao Miao, Yoshimasa Tsuruoka

机构 * The University of Tokyo(东京大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00217 2025-08-04 cs.CL cs.DB cs.LG 57%

Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges

Xiaofeng Wu, Alan Ritter, Wei Xu

机构 * College of Computing, Georgia Institute of Technology(计算学院、佐治亚理工学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00513 2025-08-04 cs.LG 50%

Text-Attributed Graph Anomaly Detection via Multi-Scale Cross- and Uni-Modal Contrastive Learning

Yiming Xu, Xu Hua, Zhen Peng, Bin Shi, Jiarun Chen, Xingbo Fu, Song Wang, Bo Dong

机构 * School of Computer Science and Technology, Xi'an Jiaotong University(西安交通大学计算机科学与技术学院) Shaanxi Provincial Key Laboratory of Big Data Knowledge Engineering, Xi’an Jiaotong University(陕西省大数据知识工程重点实验室) School of Distance Education, Xi’an Jiaotong University(西安交通大学继续教育学院) University of Virginia(弗吉尼亚大学)

专题命中 跨模态检索 :cross-modal(abstract)

Comments Accepted by ECAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 3 篇

2508.00303 2025-08-04 cs.RO 78%

TopoDiffuser: A Diffusion-Based Multimodal Trajectory Prediction Model with Topometric Maps

Zehui Xu, Junhui Wang, Yongliang Shi, Chao Gao, Guyue Zhou

机构 * School of Astronautics, Harbin Institute of Technology(哈尔滨工业大学航天学院) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) Institute of Systems Engineering and Collaborative Laboratory for Intelligent Science and Systems, Macau University of Science and Technology(澳门科学大学系统工程研究所) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院)

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18004 2025-08-04 cs.AI 57%

E.A.R.T.H.: Structuring Creative Evolution through Model Error in Generative AI

Yusen Peng, Shuhua Mao

机构 * University of Warwick(沃里克大学) Wuhan University of Technology(武汉理工大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI

Comments 44 pages,11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00583 2025-08-04 cs.NI 50%

Enhancing Wireless Networks for IoT with Large Vision Models: Foundations and Applications

Yunting Xu, Jiacheng Wang, Ruichen Zhang, Dusit Niyato, Deepu Rajan, Liang Yu, Haibo Zhou, Abbas Jamalipour, Xianbin Wang

专题命中 多模态生成 :multimodal(abstract)

Comments 7 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 9 篇

2503.06252 2025-08-04 cs.CV cs.AI 81%

Can Atomic Step Decomposition Enhance the Self-structured Reasoning of Multimodal Large Models?

Kun Xiang, Zhili Liu, Zihao Jiang, Yunshuang Nie, Kaixin Cai, Yiyang Yin, Runhui Huang, Haoxiang Fan, Hanhui Li, Weiran Huang, Yihan Zeng, Yu-Jie Yuan, Jianhua Han, Lanqing Hong, Hang Xu, Xiaodan Liang

机构 * Sun Yat-sen University(中山大学) Hong Kong University of Science and Technology(香港科学与技术大学) Shanghai Jiaotong University(上海交通大学) University of Hong Kong(香港大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2411.11930

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00726 2025-08-04 cs.CV 79%

MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models

Jiale Li, Mingrui Wu, Zixiang Jin, Hao Chen, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao, Rongrong Ji

机构 * Xiamen University(厦门大学) Zhongguancun Academy(中关村学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments ACM MM25 has accepted this paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14709 2025-08-04 cs.CV cs.RO 79%

DiFuse-Net: RGB and Dual-Pixel Depth Estimation using Window Bi-directional Parallax Attention and Cross-modal Transfer Learning

Kunal Swami, Debtanu Gupta, Amrit Kumar Muduli, Chirag Jaiswal, Pankaj Kumar Bajpai

机构 * Visual Intelligence Team, Samsung Research India Bangalore(三星印度班加罗尔视觉智能团队)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted in IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11172 2025-08-04 cs.CV 79%

TerraMesh: A Planetary Mosaic of Multimodal Earth Observation Data

Benedikt Blumenstiel, Paolo Fraccaro, Valerio Marsocci, Johannes Jakubik, Stefano Maurogiovanni, Mikolaj Czerkawski, Rocco Sedona, Gabriele Cavallaro, Thomas Brunschwiler, Juan Bernabe-Moreno, Nicolas Longépé

机构 * IBM Research – Europe(IBM欧洲研究院) European Space Agency(欧洲航天局) roman_Φ -Lab(Φ实验室) Forschungszentrum Jülich(尤利奇研究中心) University of Iceland(冰岛大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20808 2025-08-04 cs.AI 79%

MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contexts

Peijie Wang, Zhong-Zhi Li, Fei Yin, Xin Yang, Dekang Ran, Cheng-Lin Liu

机构 * MAIS, Institute of Automation of Chinese Academy of Sciences(自动化研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 45 pages, accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00615 2025-08-04 cs.LG cs.AI 57%

Similarity-Based Self-Construct Graph Model for Predicting Patient Criticalness Using Graph Neural Networks and EHR Data

Mukesh Kumar Sahu, Pinki Roy

机构 * National Institute of Technology Silchar, Assam, India(印度西里char国家理工学院)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16974 2025-08-04 cs.CV 57%

OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning

Zongyan Han, Jiale Cao, Shuo Chen, Tong Wang, Jorma Laaksonen, Rao Muhammad Anwer

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德大学人工智能学院) Tianjin University(天津大学) Nanjing University(南京大学) Aalto University(艾尔沃斯大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11849 2025-08-04 cs.CV 57%

Towards a Unified Copernicus Foundation Model for Earth Vision

Yi Wang, Zhitong Xiong, Chenying Liu, Adam J. Stewart, Thomas Dujardin, Nikolaos Ioannis Bountos, Angelos Zavras, Franziska Gerken, Ioannis Papoutsis, Laura Leal-Taixé, Xiao Xiang Zhu

机构 * Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心) National Technical University of Athens & National Observatory of Athens(雅典国家技术大学及雅典国家天文台) Harokopio University of Athens(雅典哈罗科波斯大学) NVIDIA(NVIDIA公司)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. 33 pages, 34 figures

详情

展开后加载摘要…

URL PDF HTML 收藏