arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-14 至 2025-11-14 共收录 60 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2406.16464 2025-11-14 cs.CL cs.AI cs.CV 85%

InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection

Junjie Chen, Hang Yu, Subin Huang, Sanmin Liu, Linfeng Zhang

机构 * Anhui Polytechnic University(安徽理工大学) Shanghai University(上海大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACM TOMM; Code and data are available at https://github.com/CoderChen01/InterCLIP-MEP

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10074 2025-11-14 cs.CV cs.SY eess.SY 79%

VLF-MSC: Vision-Language Feature-Based Multimodal Semantic Communication System

Gwangyeon Ahn, Jiwan Seo, Joonhyuk Kang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments To appear in the AI4NextG Workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19110 2025-11-14 cs.CV 79%

LISA: A Layer-wise Integration and Suppression Approach for Hallucination Mitigation in Multimodal Large Language Models

Zhihui Guo, Xin Man, Hui Xu, Jie Shao, Zhiguo Jiang, Xianchao Zhang, Heng Tao Shen

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10301 2025-11-14 cs.CV cs.AI 76%

Rethinking Visual Information Processing in Multimodal LLMs

Dongwan Kim, Viresh Ranjan, Takashi Nagata, Arnab Dhua, Amit Kumar K C

机构 * Seoul National University(首尔国立大学) Amazon(亚马逊)

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08238 2025-11-14 cs.CV cs.AI 73%

Remodeling Semantic Relationships in Vision-Language Fine-Tuning

Xiangyang Wu, Liu Liu, Baosheng Yu, Jiayan Qiu, Zhenwei Shi

机构 * Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院) School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) Nanyang Technological University(南洋理工大学) University of Leicester(莱斯特大学) School of Astronautics, Beihang University(北京航空航天大学航天学院)

专题命中 图文多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09973 2025-11-14 cs.CV cs.AI 62%

Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models

Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane, Naoki Makishima, Naotaka Kawata, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10098 2025-11-14 cs.CV 57%

MTAttack: Multi-Target Backdoor Attacks against Large Vision-Language Models

Zihan Wang, Guansong Pang, Wenjun Miao, Jin Zheng, Xiao Bai

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments AAAI2026, with supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09883 2025-11-14 cs.CV 57%

HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models

Liheng Zhang, Jin Wang, Hui Li, Bingfeng Zhang, Weifeng Liu

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09868 2025-11-14 cs.CV 57%

Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies

Peng Gao, Yujian Lee, Xiaofeng Zhang, Zailong Chen, Hui Zhang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 7 篇

2511.10059 2025-11-14 cs.CV 85%

When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?

Qilang Ye, Wei Zeng, Meng Liu, Jie Zhang, Yupeng Hu, Zitong Yu, Yu Zhou

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10325 2025-11-14 cs.MM 79%

TMDC: A Two-Stage Modality Denoising and Complementation Framework for Multimodal Sentiment Analysis with Missing and Noisy Modalities

Yan Zhuang, Minhao Liu, Yanru Zhang, Jiawen Deng, Fuji Ren

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05735 2025-11-14 cs.AI 79%

A Comprehensive Survey on Multi-modal Conversational Emotion Recognition with Deep Learning

Yuntao Shou, Tao Meng, Wei Ai, Fangze Fu, Nan Yin, Keqin Li

机构 * Central South University of Forestry and Technology(中部林业科技大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) State University of New York(纽约州立大学)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI

Comments 36 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09989 2025-11-14 cs.LG 78%

Towards Robust Multimodal Learning in the Open World

Fushuo Huo

机构 * Department of Computing(计算系)

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09958 2025-11-14 cs.RO cs.SD 67%

Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation

Xiangyi Wei, Haotian Zhang, Xinyi Cao, Siyu Xie, Weifeng Ge, Yang Li, Changbo Wang

机构 * School of Computer Science and Technology, East China Normal University(东华师范大学计算机科学与技术学院) School of Data Science and Engineering, East China Normal University(东华师范大学数据科学与工程学院) School of Software Engineering, East China Normal University(东华师范大学软件工程学院) School of Computer Science, Fudan University(复旦大学计算机科学学院)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08132 2025-11-14 cs.AI 57%

National Institute on Aging PREPARE Challenge: Early Detection of Cognitive Impairment Using Speech -- The SpeechCARE Solution

Maryam Zolnoori, Hossein Azadmaleki, Yasaman Haghbin, Ali Zolnour, Mohammad Javad Momeni Nezhad, Sina Rashidi, Mehdi Naserian, Elyas Esmaeili, Sepehr Karimi Arpanahi

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08670 2025-11-14 cs.HC cs.AI 57%

Once Upon an AI: Six Scaffolds for Child-AI Interaction Design, Inspired by Disney

Nomisha Kurian

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 8 篇

2511.10212 2025-11-14 cs.CV 86%

Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization

Ashutosh Anshul, Shreyas Gopal, Deepu Rajan, Eng Siong Chng

机构 * College of Computing and Data Science(计算与数据科学学院) Nanyang Technological University(南洋理工大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments Under Review, Multimodal Deepfake detection

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10134 2025-11-14 cs.CV 79%

Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction

Mingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li, Qi Zeng, Yifan Zhang, Ju Xin, Rongtao Xu, Jiguang Zhang, Xiaopeng Zhang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09773 2025-11-14 cs.LG eess.SP 78%

NeuroLingua: A Language-Inspired Hierarchical Framework for Multimodal Sleep Stage Classification Using EEG and EOG

Mahdi Samaee, Mehran Yazdi, Daniel Massicotte

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10334 2025-11-14 cs.CV 70%

Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic Alignment

Wenti Yin, Huaxin Zhang, Xiang Wang, Yuqing Lu, Yicheng Zhang, Bingquan Gong, Jialong Zuo, Li Yu, Changxin Gao, Nong Sang

专题命中 视频多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

Comments Accepted to AAAI 2026. Code is available at https://github.com/lessiYin/DSANet

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16781 2025-11-14 cs.CV cs.AI cs.CL 67%

Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features

Shihao Ji, Zihui Song

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments This paper is being withdrawn because we have identified a significant error in the implementation of our self-supervised clustering approach. Specifically, our feature aggregation step inadvertently leaked temporal information across frames, which violates the core assumption of our training-free method. We sincerely apologize to the research community

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25238 2025-11-14 cs.CV 57%

VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations

Qianqian Qiao, DanDan Zheng, Yihang Bo, Bao Peng, Heng Huang, Longteng Jiang, Huaye Wang, Jingdong Chen, Jun Zhou, Xin Jin

机构 * Nanjing University(南京大学) Huazhong University of Science and Technology(华中科技大学) Beijing Film Academy(北京电影学院) University of Science and Technology of China(中国科学技术大学) Beijing Electronic Science and Technology Institute(北京电子科技学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09870 2025-11-14 cs.CV 57%

SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object Detection

Jia Lin, Xiaofei Zhou, Jiyuan Liu, Runmin Cong, Guodao Zhang, Zhi Liu, Jiyong Zhang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to 40th AAAI Conference on Artificial Intelligence (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04369 2025-11-14 cs.CV 57%

TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding

Canhui Tang, Zifan Han, Hongbo Sun, Sanping Zhou, Xuchong Zhang, Xin Wei, Ye Yuan, Huayu Zhang, Jinglin Xu, Hao Sun

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2511.10552 2025-11-14 cs.CL 85%

URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding

Yongxin Shi, Jiapeng Wang, Zeyu Shan, Dezhi Peng, Zening Lin, Lianwen Jin

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted by AAAI 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10244 2025-11-14 cs.AI 57%

PepTriX: A Framework for Explainable Peptide Analysis through Protein Language Models

Vincent Schilling, Akshat Dubey, Georges Hattab

机构 * Center for Artificial Intelligence in Public Health Research (ZKI-PH), Robert Koch Institute(人工智能与公共健康研究所以及罗伯特·科赫研究所) Department of Mathematics and Computer Science, Free University of Berlin(数学与计算机科学系,柏林自由大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 5 篇

2511.05534 2025-11-14 cs.CL 88%

FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference

Kunxi Li, Yufan Xiong, Zhonghua Jiang, Yiyun Zhou, Zhaode Wang, Chengfei Lv, Shengyu Zhang

机构 * Zhejiang University(浙江大学) Huazhong Agricultural University(华中农业大学) Alibaba(阿里巴巴)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02034 2025-11-14 cs.CV cs.AI 73%

Abn-BLIP: Abnormality-aligned Bootstrapping Language-Image Pre-training for Pulmonary Embolism Diagnosis and Report Generation from CTPA

Zhusi Zhong, Yuli Wang, Lulu Bi, Zhuoqi Ma, Sun Ho Ahn, Christopher J. Mullin, Colin F. Greineder, Michael K. Atalay, Scott Collins, Grayson L. Baird, Cheng Ting Lin, Webster Stayman, Todd M. Kolb, Ihab Kamel, Harrison X. Bai, Zhicheng Jiao

机构 * Department of Diagnostic Imaging, Brown University Health(布朗大学健康中心诊断影像科) Warren Alpert Medical School of Brown University(布朗大学沃伦·阿尔珀特医学院) Department of Biomedical Engineering, Johns Hopkins University School of Medicine(约翰霍普金斯大学医学院生物医学工程系) Department of Radiology and Radiological Sciences, Johns Hopkins University School of Medicine(约翰霍普金斯大学医学院放射科) Johns Hopkins University Division of Pulmonary and Critical Care Medicine(约翰霍普金斯大学肺科与重症医学科) Department of Radiology, University of Colorado School of Medicine(科罗拉多大学医学院放射科)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10154 2025-11-14 cs.CV cs.AI 62%

GEA: Generation-Enhanced Alignment for Text-to-Image Person Retrieval

Hao Zou, Runqing Zhang, Xue Zhou, Jianxiao Zou

机构 * School of Automation Engineering, University of Electronic Science and Technology of China(自动化工程学院,电子科学与技术大学) Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(深圳高级研究学院,电子科学与技术大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 8pages,3figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10020 2025-11-14 cs.CV cs.AI 62%

Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation

Yuxin Jiang, Wei Luo, Hui Zhang, Qiyu Chen, Haiming Yao, Weiming Shen, Yunkang Cao

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏