arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-17 至 2025-11-17 共收录 63 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 5 篇

2511.10974 2025-11-17 cs.CV 79%

Preserving Cross-Modal Consistency for CLIP-based Class-Incremental Learning

Haoran Chen, Houze Xu, Micah Goldblum, Daoguo Dong, Zuxuan Wu

机构 * Institute of Trustworthy Embodied AI(可信具身人工智能研究院) Fudan University(复旦大学) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) Columbia University(哥伦比亚大学)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11141 2025-11-17 cs.CL cs.CY cs.LG 57%

PRSM: A Measure to Evaluate CLIP's Robustness Against Paraphrases

Udo Schlegel, Franziska Weeber, Jian Lan, Thomas Seidl

机构 * Ludwig-Maximilians-University Munich(慕尼黑路德维希-马克西米利安大学) Munich Center for Machine Learning(慕尼黑机器学习中心) University of Stuttgart(斯图加特大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

Comments 8 pages, accpeted as short paper at MMM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10769 2025-11-17 cs.CV 57%

Unifying Segment Anything in Microscopy with Vision-Language Knowledge

Manyu Li, Ruian He, Zixian Zhang, Chenxi Ma, Weimin Tan, Bo Yan

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10774 2025-11-17 cs.CV 57%

Frequency-Aware Vision-Language Multimodality Generalization Network for Remote Sensing Image Classification

Junjie Zhang, Feng Zhao, Hanqiang Liu, Jun Yu

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08172 2025-11-17 cs.AI 57%

An Efficient Training Pipeline for Reasoning Graphical User Interface Agents

Georgios Pantazopoulos, Eda B. Özyiğit

机构 * The Alan Turing Institute(艾伦·图灵研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 6 篇

2511.11124 2025-11-17 cs.CL cs.AI cs.CV cs.MM cs.SD 85%

AV-Dialog: Spoken Dialogue Models with Audio-Visual Input

Tuochao Chen, Bandhav Veluri, Hongyu Gong, Shyamnath Gollakota

机构 * Paul G. Allen School of Computer Science & Engineering, University of Washington(华盛顿大学保罗·G·阿伦计算机科学与工程学院) Meta AI Research(Meta AI研究院)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11106 2025-11-17 cs.MM cs.CV cs.SD 73%

AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache Optimization

Zhonghua Jiang, Kui Chen, Kunxi Li, Keting Yin, Yiyun Zhou, Zhaode Wang, Chengfei Lv, Shengyu Zhang

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09915 2025-11-17 cs.CL cs.MM cs.SD 73%

HI-TransPA: Hearing Impairments Translation Personal Assistant

Zhiming Ma, Shiyu Gan, Junhao Zhao, Xianming Li, Qingyun Pan, Peidong Wang, Mingjun Pan, Yuhao Mo, Jiajie Cheng, Chengxin Chen, Zhonglun Cao, Chonghan Liu, Shi Cheng

机构 * SmartFlowAI Research(SmartFlowAI研究院) Tongji University(同济大学) CMIC Guangzhou(广州CMIC) BUPT Beijing(北京邮电大学) Northeastern University Shenyang(沈阳东北大学) Peking University(北京大学) Qiyuan Tech Beijing(北京启元科技)

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10913 2025-11-17 cs.SD cs.AI cs.CR cs.MM eess.AS 67%

Synthetic Voices, Real Threats: Evaluating Large Text-to-Speech Models in Generating Harmful Audio

Guangke Chen, Yuhui Wang, Shouling Ji, Xiapu Luo, Ting Wang

机构 * Stony Brook University(石英布鲁克大学) Zhejiang University(浙江大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06606 2025-11-17 eess.AS cs.AI 62%

SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models

S Sakshi, Vaibhavi Lokegaonkar, Neil Zhang, Ramani Duraiswami, Sreyan Ghosh, Dinesh Manocha, Lie Lu

机构 * University of Maryland, College Park, USA(马里兰大学 College Park 分校) Dolby Laboratories(杜比实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments Project: https://sakshi113.github.io/spur/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27102 2025-11-17 cs.SD cs.AI eess.AS 62%

Expressive Range Characterization of Open Text-to-Audio Models

Jonathan Morse, Azadeh Naderi, Swen Gaudl, Mark Cartwright, Amy K. Hoover, Mark J. Nelson

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments Accepted at the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 9 篇

2511.11002 2025-11-17 cs.CV 79%

EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation

Zongyang Qiu, Bingyuan Wang, Xingbei Chen, Yingqing He, Zeyu Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 15 pages, 12 figures. Accepted as an Oral presentation at AAAI 2026. For code and dataset, see https://zane-zyqiu.github.io/EmoVid

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01562 2025-11-17 cs.CV 74%

Adaptive LiDAR Scanning: Harnessing Temporal Cues for Efficient 3D Object Detection via Multi-Modal Fusion

Sara Shoouri, Morteza Tavakoli Taba, Hun-Seok Kim

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

Comments Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10953 2025-11-17 cs.CV 70%

Language-Guided Graph Representation Learning for Video Summarization

Wenrui Li, Wei Han, Hengyu Man, Wangmeng Zuo, Xiaopeng Fan, Yonghong Tian

机构 * Department of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术系) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究所) School of AI for Science, the Shenzhen Graduate School, Peking University(北京大学人工智能科学学院、深圳研究生院) Peng Cheng Laboratory(鹏城实验室) School of Computer Science, Peking University(北京大学计算机学院)

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10948 2025-11-17 cs.CV cs.HC 70%

DEFT-LLM: Disentangled Expert Feature Tuning for Micro-Expression Recognition

Ren Zhang, Huilai Li, Chao qi, Guoliang Xu, Tianyu Zhou, Wei wei, Jianqin Yin

机构 * College of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications(智能工程与自动化学院,北京邮电大学)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10202 2025-11-17 cs.CL cs.CV 62%

Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval

Shubhashis Roy Dipta, Francis Ferraro

机构 * Department of Computer Science and Electrical Engineering University of Maryland Baltimore County(计算机科学与电气工程系马里兰大学巴尔的摩县)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted in IJCNLP-AACL 2025 (also presented in MAGMAR 2025 at ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12067 2025-11-17 cs.CV cs.AI cs.LG 62%

Survey of Action Recognition, Spotting and Spatio-Temporal Localization in Soccer -- Current Trends and Research Perspectives

Karolina Seweryn, Anna Wróblewska, Szymon Łukasik

机构 * NASK - National Research Institute(国家研究 institute) Faculty of Mathematics and Information Science, Warsaw University of Technology(数学与信息科学学院,华沙技术大学) Faculty of Physics and Applied Computer Science, AGH University of Science and Technology(物理与应用计算机科学学院,AGH科技大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Journal ref ACM Transactions on Intelligent Systems and Technology (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01727 2025-11-17 cs.LG cs.CV 57%

OccamVTS: Distilling Vision Models to 1% Parameters for Time Series Forecasting

Sisuo Lyu, Siru Zhong, Weilin Ruan, Qingxiang Liu, Qingsong Wen, Hui Xiong, Yuxuan Liang

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04668 2025-11-17 cs.CV 57%

SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding

Ellis Brown, Arijit Ray, Ranjay Krishna, Ross Girshick, Rob Fergus, Saining Xie

机构 * New York University(纽约大学) Boston University(波士顿大学) AllenAI Vercept

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Project page: https://ellisbrown.github.io/sims-v

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25725 2025-11-17 cs.RO 50%

A Humanoid Visual-Tactile-Action Dataset for Contact-Rich Manipulation

Eunju Kwon, Seungwon Oh, In-Chang Baek, Yucheon Park, Gyungbo Kim, JaeYoung Moon, Yunho Choi, Kyung-Joong Kim

机构 * GIST(全州科学技术大学)

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 4 篇

2511.10892 2025-11-17 cs.CV cs.AI 84%

MCN-CL: Multimodal Cross-Attention Network and Contrastive Learning for Multimodal Emotion Recognition

Feng Li, Ke Wu, Yongwei Li

机构 * School of Computer and Information Engineering, Anhui University of Finance and Economics, Anhui, 233030, China(计算机与信息工程学院,安徽财经大学,安徽,233030,中国) State Key Laboratory of Cognitive Science and Mental Health, Institute of Psychology, Chinese Academy of Sciences, Beijing, 100045, China(认知科学与心理健康国家重点实验室,心理学研究所,中国科学院,北京,100045,中国)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by 32nd International Conference on MultiMedia Modeling (MMM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11216 2025-11-17 cs.CV 83%

Positional Bias in Multimodal Embedding Models: Do They Favor the Beginning, the Middle, or the End?

Kebin Wu, Fatima Albreiki

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments accepted to AAAI 2026 main track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10997 2025-11-17 cs.CV cs.LG 83%

PROMISE: Prompt-Attentive Hierarchical Contrastive Learning for Robust Cross-Modal Representation with Missing Modalities

Jiajun Chen, Sai Cheng, Yutao Yuan, Yirui Zhang, Haitao Yuan, Peng Peng, Yi Zhong

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted by AAAI'2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09347 2025-11-17 cs.CV 57%

FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection

Jiangyong Yu, Changyong Shu, Sifan Zhou, Zichen Yu, Xing Hu, Yan Chen, Dawei Yang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments I made an operational error. I intended to update the paper with Identifier arXiv:2502.15488, not submit a new paper with a different identifier. Therefore, I would like to withdraw the current submission and resubmit an updated version for Identifier arXiv:2502.15488

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 4 篇

2511.11066 2025-11-17 cs.CV cs.AI cs.CL 83%

S2D-ALIGN: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report Generation

Jiechao Gao, Chang Liu, Yuangang Li

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07375 2025-11-17 cs.CV 79%

Novel Diffusion Models for Multimodal 3D Hand Trajectory Prediction

Junyi Ma, Wentao Bao, Jingyi Xu, Guanzhong Sun, Xieyuanli Chen, Hesheng Wang

机构 * IRMV Lab, the Department of Automation, Shanghai Jiao Tong University(IRMV实验室,自动化系,上海交通大学) Meta Reality Labs(Meta现实实验室) the Department of Electronic Engineering, Shanghai Jiao Tong University(电子工程系,上海交通大学) the School of Information and Control Engineering, China University of Mining and Technology(信息与控制工程学院,中国矿业大学) the College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11434 2025-11-17 cs.CV 70%

WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation

Wei Chow, Jiachun Pan, Yongyuan Liang, Mingze Zhou, Xue Song, Liyu Jia, Saining Zhang, Siliang Tang, Juncheng Li, Fengda Zhang, Weijia Wu, Hanwang Zhang, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) University of Maryland, College Park(马里兰大学学院公园分校) Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07334 2025-11-17 cs.CV 57%

Unleashing the Potential of Large Language Models for Text-to-Image Generation through Autoregressive Representation Alignment

Xing Xie, Jiawei Liu, Ziyue Lin, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 14 篇

2511.11407 2025-11-17 cs.CV 85%

MicroVQA++: High-Quality Microscopy Reasoning Dataset with Weakly Supervised Graphs for Multimodal Large Language Model

Manyu Li, Ruian He, Chenxi Ma, Weimin Tan, Bo Yan

机构 * Fudan University(复旦大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11503 2025-11-17 eess.SP 82%

SynthSoM-Twin: A Multi-Modal Sensing-Communication Digital-Twin Dataset for Sim2Real Transfer via Synesthesia of Machines

Junlong Chen, Ziwei Huang, Xuesong Cai, Xiang Cheng, Liuqing Yang

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏