arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-30 至 2025-07-30 共收录 52 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 8 篇

2506.18985 2025-07-30 cs.CV cs.AI 82%

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

Guanxi Shen

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments Keywords: Explainable Computer Vision, Large Vision-Language Models, AI Interpretability, Explainable AI, Visual Saliency, Attribution Maps, Cross-Modal Attribution, Human Attention Alignment, AI Transparency

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21489 2025-07-30 cs.CV 77%

Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval

Zhichuan Wang, Yang Zhou, Zhe Liu, Rui Yu, Song Bai, Yulong Wang, Xinwei He, Xiang Bai

机构 * Huazhong Agricultural University(华中农业大学) Shenzhen University(深圳大学) The University of Hong Kong(香港大学) University of Louisville(路易斯安那大学) ByteDance(字节跳动) Huazhong University of Science and Technology(华中科技大学)

专题命中 图文多模态 :multi-modal(abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22024 2025-07-30 eess.IV cs.CV 70%

Cardiac-CLIP: A Vision-Language Foundation Model for 3D Cardiac CT Images

Yutao Hu, Ying Zheng, Shumei Miao, Xiaolei Zhang, Jiahao Xia, Yaolei Qi, Yiyang Zhang, Yuting He, Qian Chen, Jing Ye, Hongyan Qiao, Xiuhua Hu, Lei Xu, Jiayin Zhang, Hui Liu, Minwen Zheng, Yining Wang, Daimin Zhang, Ji Zhang, Wenqi Shao, Yun Liu, Longjiang Zhang, Guanyu Yang

专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21246 2025-07-30 cs.CV cs.AI 62%

On Explaining Visual Captioning with Hybrid Markov Logic Networks

Monika Shah, Somdeb Sarkhel, Deepak Venugopal

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21794 2025-07-30 cs.CV 57%

Distribution-Based Masked Medical Vision-Language Model Using Structured Reports

Shreyank N Gowda, Ruichi Zhang, Xiao Gu, Ying Weng, Lu Yang

机构 * School of Computer Science, University of Nottingham(计算机科学学院,诺丁汉大学) Department of Computer Science and Technology, School of Informatics, Xiamen University(计算机科学与技术系,信息学院,厦门大学) CHI Lab, University of Oxford(CHI实验室,牛津大学) School of Computer Science, University of Nottingham Ningbo China(计算机科学学院,宁波大学中国)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted in MICCAI-W 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21291 2025-07-30 cs.CV 57%

Fairness and Robustness of CLIP-Based Models for Chest X-rays

Théo Sourget, David Restrepo, Céline Hudelot, Enzo Ferrante, Stergios Christodoulidis, Maria Vakalopoulou

机构 * MICS, CentraleSupélec - Université Paris-Saclay(MICS,中央超导学院——巴黎萨克雷大学) CONICET, Universidad de Buenos Aires(CONICET,布宜诺斯艾利斯大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted for publication at the FAIMI MICCAI workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21165 2025-07-30 eess.IV cs.CV 57%

Querying GI Endoscopy Images: A VQA Approach

Gaurav Parajuli

机构 * Johannes Kepler University Linz(约翰·凯撒大学林茨)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19651 2025-07-30 cs.CV cs.LG cs.PF 57%

PEVLM: Parallel Encoding for Vision-Language Models

Letian Kang, Shixian Luo, Yiqiang Li, Yuxin Yin, Shenxuan Zhou, Xiaoyang Yu, Jin Yang, Yong Wu

机构 * Li Auto Inc.(利自动公司)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 5 篇

2409.09545 2025-07-30 cs.SD cs.LG eess.AS 85%

Multi-Microphone and Multi-Modal Emotion Recognition in Reverberant Environment

Ohad Cohen, Gershon Hazan, Sharon Gannot

机构 * Faculty of Engineering Bar-Ilan University Ramat-Gan, Israel(工程学院 巴伊兰大学 拉马特-甘, 以色列)

专题命中 音频语音多模态 :multi-modal(title,abstract);multimodal(abstract);audio-visual(abstract);分类 eess.AS

Comments 5 pages, 4 figures, 2 tables. Accepted to EUSIPCO 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21588 2025-07-30 cs.AI cs.CV 84%

Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental Learning

Jiong Yin, Liang Li, Jiehua Zhang, Yuhan Gao, Chenggang Yan, Xichun Sheng

机构 * Hangzhou Dianzi University(杭州电子科技大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Xi’an Jiaotong University(西安交通大学) Macao Polytechnic University(澳门 polytechnic university)

专题命中 音频语音多模态 :audio-visual(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11274 2025-07-30 cs.CL cs.AI 73%

Task Arithmetic for Language Expansion in Speech Translation

Yao-Fei Cheng, Hayato Futami, Yosuke Kashiwagi, Emiru Tsunoo, Wen Shen Teo, Siddhant Arora, Shinji Watanabe

机构 * University of Washington(华盛顿大学) Sony Group Corporation(索尼集团) University of Electro-Communications(电通大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 音频语音多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14915 2025-07-30 cs.MM cs.SD eess.AS 62%

Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling

Xiaojie Li, Ronghui Li, Shukai Fang, Shuzhao Xie, Xiaoyang Guo, Jiaqing Zhou, Junkun Peng, Zhi Wang

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ByteDance Games(字节跳动游戏)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22008 2025-07-30 cs.CV 57%

VeS: Teaching Pixels to Listen Without Supervision

Sajay Raj

机构 * Indian Institute of Technology, Madras(印度理工学院马德拉斯分校)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments 6 pages, 1 figure, 1 table. Code and models are released

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 8 篇

2507.21100 2025-07-30 cs.CY cs.AI cs.CV 84%

A Tactical Behaviour Recognition Framework Based on Causal Multimodal Reasoning: A Study on Covert Audio-Video Analysis Combining GAN Structure Enhancement and Phonetic Accent Modelling

Wei Meng

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments This paper introduces a structurally innovative and mathematically rigorous framework for multimodal tactical reasoning, offering a significant advance in causal inference and graph-based threat recognition under noisy conditions

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21945 2025-07-30 cs.CV 83%

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment

Xin Wang, Peng-Jie Li, Yuan-Yuan Shen

机构 * School of Sport Engineering, Beijing Sport University(体育工程学院,北京体育大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to Applied Soft Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21161 2025-07-30 cs.CV cs.AI cs.LG 81%

Seeing Beyond Frames: Zero-Shot Pedestrian Intention Prediction with Raw Temporal Video and Multimodal Cues

Pallavi Zambare, Venkata Nikhil Thanikella, Ying Liu

机构 * Departmrnt of computer science(计算机科学系) Texas Tech University(得克萨斯科技大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in IEEE 3rd International Conference on Artificial Intelligence, Blockchain, and Internet of Things (AIBThings 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21649 2025-07-30 cs.CV 79%

The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM

Shibo Gao, Peipei Yang, Haiyang Guo, Yangyang Liu, Yi Chen, Shuai Li, Han Zhu, Jian Xu, Xu-Yao Zhang, Linlin Huang

机构 * Beijing Jiaotong University(北京交通大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)

专题命中 视频多模态 :MLLM(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10541 2025-07-30 cs.IR cs.AI 79%

Multi-Modal Hypergraph Enhanced LLM Learning for Recommendation

Xu Guo, Tong Zhang, Yuanzhi Wang, Chenxu Wang, Fuyun Wang, Xudong Wang, Xiaoya Zhang, Xin Liu, Zhen Cui

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) Shituoyun (Nanjing) Technology Co., Ltd(石图云(南京)科技有限公司) School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments 12 pages, 4 figures, submitted to IEEE Transactions on Knowledge and Data Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20177 2025-07-30 cs.CV cs.MM 73%

Towards Universal Modal Tracking with Online Dense Temporal Token Learning

Yaozong Zheng, Bineng Zhong, Qihua Liang, Shengping Zhang, Guorong Li, Xianxian Li, Rongrong Ji

机构 * Key Laboratory of Education Blockchain and Intelligent Technology, Ministry of Education(教育区块链与智能技术重点实验室,教育部) Guangxi Key Laboratory of Multi-Source Information Mining and Security, Guangxi Normal University(广西多源信息挖掘与安全重点实验室,广西师范大学) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) University of Chinese Academy of Sciences(中国科学院大学) Media Analytics and Computing Lab, Department of Artificial Intelligence, School of Informatics, Xiamen University(媒体分析与计算实验室,人工智能系,厦门大学)

专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments arXiv admin note: text overlap with arXiv:2401.01686

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21063 2025-07-30 q-bio.NC cs.CY 71%

Make Silence Speak for Itself: a multi-modal learning analytic approach with neurophysiological data

Mingxuan Gao, Jingjing Chen, Yun Long, Xiaomeng Xu, Yu Zhang

专题命中 视频多模态 :multi-modal(title)

Comments 25 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21971 2025-07-30 cs.CV 57%

EIFNet: Leveraging Event-Image Fusion for Robust Semantic Segmentation

Zhijiang Li, Haoran He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2507.21917 2025-07-30 cs.CV 83%

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

Nicola Fanelli, Gennaro Vessio, Giovanna Castellano

机构 * Department of Computer Science University of Bari Aldo Moro(计算机科学系巴里大学Aldo Moro)

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21357 2025-07-30 cs.LG 50%

A Contrastive Diffusion-based Network (CDNet) for Time Series Classification

Yaoyu Zhang, Chi-Guhn Lee

机构 * Department of Mechanical and Industrial Engineering University of Toronto(机械与工业工程系多伦多大学)

专题命中 跨模态检索 :multimodal(abstract)

Comments 19 pages, conference

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 7 篇

2507.21741 2025-07-30 cs.CV cs.MM 84%

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

Shaojun E, Yuchen Yang, Jiaheng Wu, Yan Zhang, Tiejun Zhao, Ziyan Chen

机构 * Global Tone Communication Technology Co., Ltd.(全球 tone 通信技术有限公司) Faculty of computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院)

专题命中 多模态生成 :multimodal(title,abstract);any-to-any(abstract);分类 cs.CV、cs.MM

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21430 2025-07-30 cs.AR 78%

Automated HEMT Model Construction from Datasheets via Multi-Modal Intelligence and Prior-Knowledge-Free Optimization

Yuang Peng, Jiarui Zhong, Yang Zhang, Hong Cai Chen

专题命中 多模态生成 :multi-modal(title,abstract)

Comments 12 pages, 12 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21260 2025-07-30 cs.LG cs.AI q-bio.QM 74%

Adaptive Multimodal Protein Plug-and-Play with Diffusion-Based Priors

Amartya Banerjee, Xingyu Xu, Caroline Moosmüller, Harlin Lee

机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态生成 :multimodal(title);分类 cs.AI

Comments Code: https://github.com/amartya21/Adam-PnP

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20987 2025-07-30 cs.CV cs.AI 62%

JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1

Xinhan Di, Kristin Qi, Pengqian Yu

机构 * Computer Science, University of Massachusetts Boston(马萨诸塞大学波士顿分校计算机科学系) National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments WiCV @ ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20536 2025-07-30 cs.CV cs.AI cs.HC 62%

T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation

Chieh-Yun Chen, Min Shi, Gong Zhang, Humphrey Shi

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21627 2025-07-30 cs.CV 57%

GuidPaint: Class-Guided Image Inpainting with Diffusion Models

Qimin Wang, Xinda Liu, Guohua Geng

机构 * Northwest University(西北大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21493 2025-07-30 cs.GR 50%

BANG: Dividing 3D Assets via Generative Exploded Dynamics

Longwen Zhang, Qixuan Zhang, Haoran Jiang, Yinuo Bai, Wei Yang, Lan Xu, Jingyi Yu

专题命中 多模态生成 :multimodal(abstract)

Comments Homepage: https://sites.google.com/view/bang7355608

详情

展开后加载摘要…

URL PDF HTML 收藏