arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-19 至 2025-11-19 共收录 56 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 5 篇

2508.05430 2025-11-19 cs.CV cs.AI cs.LG 79%

Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions

Hubert Baniecki, Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer, Eyke Hüllermeier, Przemyslaw Biecek

机构 * University of Warsaw(华沙大学) Warsaw University of Technology(华沙理工大学) LMU Munich(慕尼黑大学) MCML DFKI(德累斯顿大学) Bielefeld University(比勒菲尔德大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025. Code: https://github.com/hbaniecki/fixlip

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06009 2025-11-19 cs.CV 74%

Continual Learning for Image Captioning through Improved Image-Text Alignment

Bertram Taetz, Gal Bordelius

机构 * IT & Engineering International University of Applied Sciences(IT与工程国际应用科学大学)

专题命中 图文多模态 :image-text(title);分类 cs.CV

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13876 2025-11-19 cs.CV 70%

QwenCLIP: Boosting Medical Vision-Language Pretraining via LLM Embeddings and Prompt tuning

Xiaoyang Wei, Camille Kurtz, Florence Cloppet

机构 * Laboratoire d'Informatique Paris Descartes (LIPADE), Université Paris Cité (France)(巴黎笛卡尔大学信息学实验室(LIPADE),巴黎城市大学(法国))

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments This work has been submitted to the IEEE ISBI for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07463 2025-11-19 cs.RO cs.AI cs.CV 62%

DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving

Sven Kirchner, Nils Purschke, Ross Greer, Alois C. Knoll

机构 * Chair of Robotics, Artificial Intelligence and Real-time Systems, Technical University of Munich(机器人学、人工智能与实时系统教授会,慕尼黑技术大学) Computer Science and Engineering Department, University of California Merced(计算机科学与工程系,加州大学默塞德分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13881 2025-11-19 cs.CV 57%

VLMs Guided Interpretable Decision Making for Autonomous Driving

Xin Hu, Taotao Jing, Renran Tian, Zhengming Ding

机构 * Department of Computer Science, Tulane University(路易斯安那大学计算机科学系) Qualcomm(高通公司) Department of Industrial and Systems Engineering, North Carolina State University(北卡罗来纳州立大学工业与系统工程系)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 5 篇

2408.14595 2025-11-19 cs.CL 88%

Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models

Ian Stewart, Sameera Horawalavithana, Brendan Kennedy, Sai Munikoti, Karl Pazdernik

机构 * Pacific Northwest National Laboratory(太平洋西北国家实验室)

专题命中 音频语音多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CL

Comments arxiv

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01711 2025-11-19 cs.CV cs.AI 84%

GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval

Bowen Yang, Yun Cao, Chen He, Xiaosu Su

机构 * School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05817 2025-11-19 cs.HC cs.MM cs.SD 79%

TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech

Weiyan Shi, Sunaya Upadhyay, Geraldine Quek, Kenny Tsu Wei Choo

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments Accepted at AAAI 2026 Workshop on Creative AI for Live Interactive Performances (CLIP). To be published in Springer CCIS series

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14255 2025-11-19 cs.CL 57%

AfriSpeech-MultiBench: A Verticalized Multidomain Multicountry Benchmark Suite for African Accented English ASR

Gabrial Zencha Ashungafac, Mardhiyah Sanni, Busayo Awobade, Alex Gichamba, Tobi Olatunji

机构 * Intron Health(Intron健康)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted As a Conference Paper IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10596 2025-11-19 cs.CY cs.AI cs.HC 57%

GenAI Voice Mode in Programming Education

Sven Jacobs, Natalie Kiesler

机构 * University of Siegen(锡根大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted for the 25th International Conference on Computing Education Research (Koli Calling '25)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 7 篇

2511.14057 2025-11-19 cs.LG cs.AI 79%

A Machine Learning-Based Multimodal Framework for Wearable Sensor-Based Archery Action Recognition and Stress Estimation

Xianghe Liu, Jiajia Liu, Chuxian Xu, Minghan Wang, Hongbo Peng, Tao Sun, Jiaqi Xu

机构 * Beijing PsychTech Technology Co., Ltd.(北京心理科技技术有限公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14186 2025-11-19 cs.CV cs.AI 62%

Few-Shot Precise Event Spotting via Unified Multi-Entity Graph and Distillation

Zhaoyu Liu, Kan Jiang, Murong Ma, Zhe Hou, Yun Lin, Jin Song Dong

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments The 40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14119 2025-11-19 cs.MM cs.AI 62%

Real-Time Mobile Video Analytics for Pre-arrival Emergency Medical Services

Liuyi Jin, Amran Haroon, Radu Stoleru, Pasan Gunawardena, Michael Middleton, Jeeeun Kim

机构 * Engineering, Texas A\&M University Emergency Medical Services (EMS), Texas A\&M University

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14249 2025-11-19 cs.CL 57%

Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction Learning

Rui Liu, Yuan Zhao, Zhenqi Jia

机构 * Rui Liu \equalcontrib , Yuan Zhao \equalcontrib , Zhenqi Jia(Rui Liu、Yuan Zhao、Zhenqi Jia)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13993 2025-11-19 cs.CV 57%

Learning Skill-Attributes for Transferable Assessment in Video

Kumar Ashutosh, Kristen Grauman

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025, Project webpage: https://vision.cs.utexas.edu/projects/CrossTrainer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16597 2025-11-19 cs.CV 57%

EventHallusion: Diagnosing Event Hallucinations in Video LLMs

Jiacheng Zhang, Yang Jiao, Shaoxiang Chen, Na Zhao, Zhiyu Tan, Hao Li, Xingjun Ma, Jingjing Chen

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14661 2025-11-19 cs.HC 50%

M-CALLM: Multi-level Context Aware LLM Framework for Group Interaction Prediction

Diana Romero, Xin Gao, Daniel Khalkhali, Salma Elmalaki

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 4 篇

2511.11305 2025-11-19 cs.IR cs.AI cs.CV cs.LG 81%

MOON Embedding: Multimodal Representation Learning for E-commerce Search Advertising

Chenghan Fu, Daoze Zhang, Yukang Lin, Zhanheng Nie, Xiang Zhang, Jianyu Liu, Yueran Liu, Wanxian Guan, Pengjie Wang, Jian Xu, Bo Zheng

机构 * Alibaba Group(阿里巴巴集团) Taobao(淘宝)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 31 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14416 2025-11-19 cs.LG 78%

Toward Robust and Harmonious Adaptation for Cross-modal Retrieval

Haobin Li, Mouxing Yang, Xi Peng

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) National Key Laboratory of Fundamental Algorithms and Models for Engineering Numerical Simulation, Sichuan University(四川省工程数值模拟基础算法与模型国家重点实验室)

专题命中 跨模态检索 :cross-modal(title,abstract)

Comments 19 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14229 2025-11-19 cs.LG 67%

EBind: a practical approach to space binding

Jim Broadbent, Felix Cohen, Frederik Hvilshøj, Eric Landau, Eren Sasoglu

机构 * Encord London, UK(Encord伦敦)

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03989 2025-11-19 cs.LG cs.AI 57%

Dynamic User-controllable Privacy-preserving Few-shot Sensing Framework

Ajesh Koyatan Chathoth, Shuhao Yu, Stephen Lee

机构 * University of Pittsburgh(匹兹堡大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 8 篇

2511.09611 2025-11-19 cs.CV 83%

MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation

Ye Tian, Ling Yang, Jiongfan Yang, Anran Wang, Yu Tian, Jiani Zheng, Haochen Wang, Zhiyang Teng, Zhuochen Wang, Yinjie Wang, Yunhai Tong, Mengdi Wang, Xiangtai Li

机构 * Peking University(北京大学) ByteDance(字节跳动) Princeton University(普林斯顿大学) CASIA(中国科学院自动化研究所) The University of Chicago(芝加哥大学)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Project Page: https://tyfeld.github.io/mmadaparellel.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14027 2025-11-19 cs.CL 77%

HiEAG: Evidence-Augmented Generation for Out-of-Context Misinformation Detection

Junjie Wu, Yumeng Fu, Nan Yu, Guohong Fu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Institute of Artificial Intelligence, Soochow University(苏州大学人工智能研究院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13689 2025-11-19 cs.CL cs.CV 76%

Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation

Sofia Jamil, Kotla Sai Charan, Sriparna Saha, Koustava Goswami, Joseph K J

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13442 2025-11-19 cs.CV cs.AI 73%

Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline

Rui Zuo, Qinyue Tong, Zhe-Ming Lu, Ziqian Lu

机构 * Zhejiang University(浙江大学) Zhejiang Sci-Tech University(浙江科技学院)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14760 2025-11-19 cs.CV 70%

UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning

Rui Tian, Mingfei Gao, Haiming Gang, Jiasen Lu, Zhe Gan, Yinfei Yang, Zuxuan Wu, Afshin Dehghan

机构 * Institute of Trustworthy Embodied AI, Fudan University(可信具身人工智能研究院,复旦大学) Apple(苹果公司)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01119 2025-11-19 cs.CV cs.LG 70%

The Promise of RL for Autoregressive Image Editing

Saba Ahmadi, Rabiul Awal, Ankur Sikarwar, Amirhossein Kazemnejad, Ge Ya Luo, Juan A. Rodriguez, Sai Rajeswar, Siva Reddy, Christopher Pal, Benno Krojer, Aishwarya Agrawal

机构 * Mila – Quebec AI Institute(魁北克AI研究所) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) École de Technologie Supérieure (ETS)(高等技术学院) Polytechnique Montréal(蒙特利尔理工学院) ServiceNow(ServiceNow公司) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.20105 2025-11-19 cs.CV 57%

FreeSeg-Diff: Training-Free Open-Vocabulary Segmentation with Diffusion Models

Barbara Toniella Corradini, Mustafa Shukor, Paul Couairon, Guillaume Couairon, Franco Scarselli, Matthieu Cord

机构 * DIISM University of Siena(DIISM锡耶纳大学) CNRS, ISIR Sorbonne University(CNRS,ISIR索邦大学) Inria, ARCHES Sorbonne University(Inria,ARCHES索邦大学) Valeo.ai Sorbonne University(Valeo.ai索邦大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Journal ref Proceedings of the 2025 International Joint Conference on Neural Networks (IJCNN 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14453 2025-11-19 q-bio.NC 50%

Multi-network Topology Underlying Individual Language Learning Success

Peilun Song, Shuguang Yang, Xiujuan Geng, Zhenzhong Gan, Suiping Wang, Gangyi Feng

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 12 篇

2511.13135 2025-11-19 cs.CV 85%

MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation

Junjie Yang, Yuhao Yan, Gang Wu, Yuxuan Wang, Ruoyu Liang, Xinjie Jiang, Xiang Wan, Fenglei Fan, Yongquan Zhang, Feiwei Qin, Changmiao Wang

机构 * South China University of Technology(华南理工大学) Sun Yat-sen University(中山大学) Hangzhou Dianzi University(杭州电子科技大学) Zhejiang University of Finance & Economics(浙江财经大学) National University of Singapore(新加坡国立大学) Shenzhen Research Institute of Big Data(深圳大数据研究院) City University of Hong Kong(香港城市大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments CVPR 2026 Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏