arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-19 至 2025-08-19 共收录 62 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2507.19196 2025-08-19 cs.RO cs.CL cs.HC 80%

Towards Multimodal Social Conversations with Robots: Using Vision-Language Models

Ruben Janssens, Tony Belpaeme

机构 * Ghent University–imec(根特大学–imec)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at the workshop "Human - Foundation Models Interaction: A Focus On Multimodal Information" (FoMo-HRI) at IEEE RO-MAN 2025 (Camera-ready version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12458 2025-08-19 cs.CL 79%

M3PO: Multimodal-Model-Guided Preference Optimization for Visual Instruction Following

Ruirui Gao, Emily Johnson, Bowen Tan, Yanfei Qian

机构 * University of Massachusetts, Amherst(马萨诸塞大学阿姆赫斯特分校)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12198 2025-08-19 physics.ao-ph cs.AI cs.LG 79%

Exploring Multimodal AI Reasoning for Meteorological Forecasting from Skew-T Diagrams

ChangJae Lee, Heecheol Yang, Jonghak Choi

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments 24 pages, 3 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12137 2025-08-19 cs.CV 79%

Infusing fine-grained visual knowledge to Vision-Language Models

Nikolaos-Antonios Ypsilantis, Kaifeng Chen, André Araujo, Ondřej Chum

机构 * VRG, FEE, Czech Technical University in Prague(捷克布拉格技术大学)

专题命中 图文多模态 :multimodal(abstract,comments);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments ICCVW 2025 accepted paper. Workshop name: "What is Next in Multimodal Foundation Models?"

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.01955 2025-08-19 cs.CV 74%

Adaptively Clustering Neighbor Elements for Image-Text Generation

Zihua Wang, Xu Yang, Hanwang Zhang, Haiyang Xu, Ming Yan, Fei Huang, Yu Zhang

机构 * School of Computer Science and Engineering, Nanyang Technological University(计算机科学与工程学院,南洋理工大学) DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

专题命中 图文多模态 :image-text(title);分类 cs.CV

Comments This work has been accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12109 2025-08-19 cs.CV cs.AI 62%

Simple o3: Towards Interleaved Vision-Language Reasoning

Ye Wang, Qianglong Chen, Zejun Li, Siyuan Wang, Shijie Guo, Zhirui Zhang, Zhongyu Wei

机构 * Independent Researcher(独立研究者)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12455 2025-08-19 cs.CV 57%

X-Ray-CoT: Interpretable Chest X-ray Diagnosis with Vision-Language Models via Chain-of-Thought Reasoning

Chee Ng, Liliang Sun, Shaoqing Tang

机构 * Universiti Teknologi Malaysia(马来西亚技术大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12400 2025-08-19 cs.CV 57%

MPCAR: Multi-Perspective Contextual Augmentation for Enhanced Visual Reasoning in Large Vision-Language Models

Amirul Rahman, Qiang Xu, Xueying Huang

机构 * University of Malaya(马来大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10357 2025-08-19 cs.CV 57%

Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models

Enming Zhang, Bingke Zhu, Yingying Chen, Qinghai Miao, Ming Tang, Jinqiao Wang

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) Wuhan AI Research(武汉人工智能研究所) Peng Cheng Laboratory(鹏城实验室)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 6 篇

2508.12842 2025-08-19 cs.CV cs.MM 88%

Multi-source Multimodal Progressive Domain Adaption for Audio-Visual Deception Detection

Ronghao Lin, Sijie Mai, Ying Zeng, Qiaolin He, Aolin Xiong, Haifeng Hu

机构 * Sun Yat-Sen University(中山大学) Nanyang Technological University(南洋理工大学) South China Normal University(华南师范大学)

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(title,abstract);分类 cs.CV、cs.MM

Comments Accepted at ACM MM 2025 SVC Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05934 2025-08-19 cs.CR cs.AI 83%

Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models

Ma Teng, Jia Xiaojun, Duan Ranjie, Li Xinfeng, Huang Yihao, Jia Xiaoshuang, Chu Zhixuan, Ren Wenqi

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21847 2025-08-19 cs.CV cs.SD 70%

Differentiable Room Acoustic Rendering with Multi-View Vision Priors

Derong Jin, Ruohan Gao

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.CV

Comments ICCV 2025 (Oral); Project Page: https://humathe.github.io/avdar/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24115 2025-08-19 cs.CL cs.MM 62%

TeleAntiFraud-28k: An Audio-Text Slow-Thinking Dataset for Telecom Fraud Detection

Zhiming Ma, Peidong Wang, Minhua Huang, Jingpeng Wang, Kai Wu, Xiangzhao Lv, Yachun Pang, Yin Yang, Wenjie Tang, Yuchen Kang

机构 * China Mobile Internet Company Ltd.(中国移动互联网有限公司) Northeastern University(东北大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12666 2025-08-19 eess.AS 57%

Cryfish: On deep audio analysis with Large Language Models

Anton Mitrofanov, Sergei Novoselov, Tatiana Prisyach, Vladislav Marchevskiy, Arseniy Karelin, Nikita Khmelev, Dmitry Dutov, Stepan Malykh, Igor Agafonov, Aleksandr Nikitin, Oleg Petrov

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Journal ref Proc. Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12186 2025-08-19 cs.SI 50%

MAD: A Benchmark for Multi-Turn Audio Dialogue Fact-Checking

Chaewan Chun, Lysandre Terrisse, Delvin Ce Zhang, Dongwon Lee

专题命中 音频语音多模态 :multimodal(abstract)

Comments 11 pages, Accepted to SBP-BRiMS 2025 Working Paper

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2507.14784 2025-08-19 cs.CV cs.AI 73%

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering

Xinxin Dong, Baoyun Peng, Haokai Ma, Yufei Wang, Zixuan Dong, Fei Hu, Xiaodong Wang

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12586 2025-08-19 cs.CV 70%

Foundation Model for Skeleton-Based Human Action Understanding

Hongsong Wang, Wanjiang Weng, Junbo Wang, Fang Zhao, Guo-Sen Xie, Xin Geng, Liang Wang

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及交叉应用关键实验室(东南大学),中华人民共和国教育部,中国) School of Software, Northwestern Polytechnical University(西北工业大学软件学院) State Key Laboratory for Novel Software Technology and School of Intelligence Science and Technology, Nanjing University(新型软件技术国家重点实验室和南京大学智能科学与技术学院) School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by TPAMI, Code is available at: https://github.com/wengwanjiang/FoundSkelModel

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11903 2025-08-19 cs.CV 70%

OVG-HQ: Online Video Grounding with Hybrid-modal Queries

Runhao Zeng, Jiaqi Mao, Minghao Lai, Minh Hieu Phan, Yanjie Dong, Wei Wang, Qi Chen, Xiping Hu

机构 * Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(人工智能研究院,深圳MSU-BIT大学) University of Adelaide(阿德莱德大学)

专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24039 2025-08-19 cs.CV cs.HC 57%

Foundation Models for Zero-Shot Segmentation of Scientific Images without AI-Ready Data

Shubhabrata Mukherjee, Jack Lang, Obeen Kwon, Iryna Zenyuk, Valerie Brogden, Adam Weber, Daniela Ushizima

机构 * Lawrence Berkeley National Laboratory(伯克利国家实验室) University of California, Irvine(加州大学尔湾分校) University of California, Berkeley(加州大学伯克利分校) Covalent Metrology(协力计量)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments This paper has been accepted for presentation at the 59th International Conference on Parallel Processing (ICPP 2025), DRAI workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02356 2025-08-19 cs.CV 57%

InterRVOS: Interaction-aware Referring Video Object Segmentation

Woojeong Jin, Seongchan Kim, Jaeho Lee, Seungryong Kim

专题命中 视频多模态 :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 4 篇

2401.10011 2025-08-19 cs.CV 85%

CPCL: Cross-Modal Prototypical Contrastive Learning for Weakly Supervised Text-based Person Retrieval

Xinpeng Zhao, Yanwei Zheng, Chuanlin Lan, Xiaowei Zhang, Bowen Huang, Jibin Yang, Dongxiao Yu

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);image-text(abstract);分类 cs.CV

Comments 9 pages, 6 figures, under peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11673 2025-08-19 cs.LG cs.AI cs.CV cs.MM 82%

Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning

Haojie Zhang, Yixiong Liang, Hulin Kuang, Lihui Cen, Zhe Qu, Yigang Cen, Min Zeng, Shichao Kan

机构 * School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学) School of Automation, Central South University(自动化学院,中南大学) School of Computer Science and Technology, Beijing Jiaotong University(计算机科学与技术学院,北京交通大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments 10 pages, 3 figures, submitted to ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12149 2025-08-19 cs.AI 79%

MOVER: Multimodal Optimal Transport with Volume-based Embedding Regularization

Haochen You, Baojing Liu

机构 * Columbia University(哥伦比亚大学) Hebei Institute of Communications(河北通信学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Accepted as a conference paper at CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12108 2025-08-19 cs.CV 57%

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine

Ziyang Zhang, Yang Yu, Xulei Yang, Si Yong Yeo

机构 * MedVisAI Lab Department of ECE Northwestern University(MedVisAI实验室 电子工程系 西北大学) Institute for Infocomm Research (I 2 R) A*STAR, Singapore(信息与通信研究所(I 2 R)A*STAR,新加坡) MedVisAI Lab Lee Kong Chian School of Medicine, Nanyang Technological University(MedVisAI实验室 李科田医学院,南洋理工大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 9 篇

2508.12854 2025-08-19 cs.AI cs.CL cs.CV cs.HC cs.MM 83%

E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model

Ronghao Lin, Shuai Shen, Weipeng Hu, Qiaolin He, Aolin Xiong, Li Huang, Haifeng Hu, Yap-peng Tan

机构 * Sun Yat-sen University(中山大学) Nanyang Technological University(南洋理工大学) Desay SV Automotive Co., Ltd(德赛西威汽车有限公司) Pazhou Laboratory(琶洲实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at ACM MM 2025 Grand Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12399 2025-08-19 cs.CV 79%

Federated Cross-Modal Style-Aware Prompt Generation

Suraj Prasad, Navyansh Mahla, Sunny Gupta, Amit Sethi

机构 * Indian Institute of Technology Bombay(印度理工学院孟买学院)

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10118 2025-08-19 cs.LG cs.CV 79%

From Intent to Execution: Multimodal Chain-of-Thought Reinforcement Learning for Precise CAD Code Generation

Ke Niu, Haiyang Yu, Zhuofan Chen, Mengyang Zhao, Teng Fu, Bin Li, Xiangyang Xue

机构 * Fudan University(复旦大学) ByteDance Inc(字节跳动公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12512 2025-08-19 cs.CV 57%

LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models

Krishna Teja Chitty-Venkata, Murali Emani, Venkatram Vishwanath

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICIP 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12396 2025-08-19 cs.CV 57%

DeCoT: Decomposing Complex Instructions for Enhanced Text-to-Image Generation with Large Language Models

Xiaochuan Lin, Xiangyong Chen, Xuan Li, Yichen Su

机构 * Henan Polytechnic University(河南理工大学)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12361 2025-08-19 cs.LG cs.AI math.ST stat.TH 57%

Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models

Xun Su, Jianming Huang, Yang Yusen, Zhongxi Fang, Hiroyuki Kasai

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏