arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-07 至 2025-08-07 共收录 77 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2508.04175 2025-08-07 cs.CV 83%

AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward Optimization

Jingyi Liao, Yongyi Su, Rong-Cheng Tu, Zhao Jin, Wenhao Sun, Yiting Li, Dacheng Tao, Xun Xu, Xulei Yang

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15807 2025-08-07 cs.CV cs.AI 81%

True Multimodal In-Context Learning Needs Attention to the Visual Context

Shuo Chen, Jianzhe Liu, Zhen Han, Yan Xia, Daniel Cremers, Philip Torr, Volker Tresp, Jindong Gu

机构 * LMU Munich(慕尼黑大学) Technical University of Munich(慕尼黑技术大学) Siemens AG(西门子股份公司) University of Science and Technology of China(中国科学技术大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Konrad Zuse School of Excellence in Reliable AI (relAI)(Konrad Zuse可靠性人工智能卓越学院) University of Oxford(牛津大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04028 2025-08-07 cs.CV cs.IR 79%

Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval

Yifan Wang, Tao Wang, Chenwei Tang, Caiyang Yu, Zhengqing Zang, Mengmi Zhang, Shudong Huang, Jiancheng Lv

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) Engineering Research Center of Machine Learning and Industry Intelligence, Ministry of Education, Chengdu, China(教育部机器学习与产业智能工程研究中心) Deep NeuroCognition Lab, I2R and CFAR, Agency for Science, Technology and Research, Singapore(深度神经认知实验室,I2R和CFAR,科技研究局,新加坡)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments 10 pages, 7figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04469 2025-08-07 cs.CV cs.CL 62%

FrEVL: Leveraging Frozen Pretrained Embeddings for Efficient Vision-Language Understanding

Emmanuelle Bourigault, Pauline Bourigault

机构 * Department of Engineering Science, University of Oxford(牛津大学工程科学系) Department of Electrical Engineering, Imperial College London(伦敦帝国学院电子工程系)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01697 2025-08-07 cs.CV cs.AI 62%

Hulk: A Universal Knowledge Translator for Human-Centric Tasks

Yizhou Wang, Yixuan Wu, Weizhen He, Xun Guo, Feng Zhu, Lei Bai, Rui Zhao, Jian Wu, Tong He, Wanli Ouyang, Shixiang Tang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by TPAMI2025

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, Jul. 2025, pp. 5672-5689, vol. 47

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02762 2025-08-07 cs.LG cs.AI 57%

Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment

Dahun Kim, Anelia Angelova

机构 * Google DeepMind(谷歌DeepMind)

专题命中 图文多模态 :image-text(abstract);分类 cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 12 篇

2508.04566 2025-08-07 cs.CV cs.AI cs.MM 90%

CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization

Jinxing Zhou, Ziheng Zhou, Yanghao Zhou, Yuxin Mao, Zhangling Duan, Dan Guo

专题命中 音频语音多模态 :cross-modal(title,abstract);audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06362 2025-08-07 cs.CV cs.MM cs.SD eess.AS 89%

Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs

Umberto Cappellazzo, Minsu Kim, Stavros Petridis

机构 * Imperial College London(伦敦帝国学院)

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted to IEEE ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04418 2025-08-07 cs.MM cs.CV cs.MA cs.SD eess.AS 85%

Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation

Jinxing Zhou, Yanghao Zhou, Mingfei Han, Tong Wang, Xiaojun Chang, Hisham Cholakkal, Rao Muhammad Anwer

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Project page: https://github.com/jasongief/TGS-Agent

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04129 2025-08-07 cs.CV 83%

SVC 2025: the First Multimodal Deception Detection Challenge

Xun Lin, Xiaobao Guo, Taorui Wang, Yingjie Ma, Jiajian Huang, Jiayu Zhang, Junzhe Cao, Zitong Yu

机构 * Great Bay University(大亚湾大学) Nanyang Technological University(南洋理工大学)

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV

Comments Accepted by Workshop SVC of ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03724 2025-08-07 cs.CV 83%

From Waveforms to Pixels: A Survey on Audio-Visual Segmentation

Jia Li, Yapeng Tian

机构 * Department of Computer Science, The University of Texas at Dallas(计算机科学系,德克萨斯大学达拉斯分校)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04353 2025-08-07 cs.MM cs.AI 81%

LUST: A Multi-Modal Framework with Hierarchical LLM-based Scoring for Learned Thematic Significance Tracking in Multimedia Content

Anderson de Lima Luiz

机构 * AImotion Bavaria Technische Hochschule Ingolstadt(巴伐利亚AImotion技术大学)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI、cs.MM

Comments 5 pages and 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15641 2025-08-07 cs.CL cs.AI 81%

Leveraging Context for Multimodal Fallacy Classification in Political Debates

Alessio Pittiglio

机构 * DISI, University of Bologna(DISI,博洛尼亚大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 12th Workshop on Argument Mining (ArgMining 2025) @ ACL 2025

Journal ref In Proceedings of the 12th Argument mining Workshop (ArgMining 2025), pages 388-397, Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02786 2025-08-07 cs.SD cs.CV eess.AS 81%

CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation

Yuanhong Chen, Kazuki Shimada, Christian Simon, Yukara Ikemiya, Takashi Shibuya, Yuki Mitsufuji

机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) Sony Group Corporation(索尼集团) Sony AI, Sony Group Corporation(索尼人工智能,索尼集团)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07384 2025-08-07 cs.SD eess.AS 79%

AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment

Yu Chen, Hongxu Zhu, Jiadong Wang, Kainan Chen, Xinyuan Qian

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10416 2025-08-07 cs.MM cs.SD eess.AS 73%

Can Sound Replace Vision in LLaVA With Token Substitution?

Ali Vosoughi, Jing Bi, Pinxin Liu, Yunlong Tang, Chenliang Xu

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.MM、eess.AS

Comments Project page: https://ali-vosoughi.github.io/SoundCLIP/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02929 2025-08-07 cs.HC cs.SD eess.AS 70%

AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design in Augmented Reality

Brandon Woodard, Margarita Geleta, Joseph J. LaViola, Andrea Fanelli, Rhonda Wilson

机构 * Brown University(布朗大学) University of California at Berkeley(加州大学伯克利分校) University of Central Florida(中央佛罗里达大学) Dolby Laboratories(杜比实验室)

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 eess.AS

Comments Updated abstract

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04481 2025-08-07 cs.LG cs.HC cs.NE cs.SD eess.AS 57%

Emotion Detection Using Conditional Generative Adversarial Networks (cGAN): A Deep Learning Approach

Anushka Srivastava

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 3 pages, 2 tables, submitted for arXiv preprint

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 1 篇

2508.03722 2025-08-07 cs.CV cs.AI 86%

Multimodal Video Emotion Recognition with Reliable Reasoning Priors

Zhepeng Wang, Yingjian Zhu, Guanghao Dong, Hongzhu Yi, Feng Chen, Xinming Wang, Jun Xie

机构 * Lenovo Research(联想研究院) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) Institute of Automation, CAS(中国科学院自动化研究所) Macau University of Science and Technology(澳门科学理工学院) School of Computer Science and Technology, UCAS(中国科学院大学计算机科学与技术学院)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 3 篇

2507.19264 2025-08-07 cs.CV 79%

SimMLM: A Simple Framework for Multi-modal Learning with Missing Modality

Sijie Li, Chen Chen, Jungong Han

机构 * School of Computer Science, University of Sheffield, UK(计算机科学学院,谢菲尔德大学)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03729 2025-08-07 cs.LG cs.HC cs.MM 74%

Privileged Contrastive Pretraining for Multimodal Affect Modelling

Kosmas Pinitas, Konstantinos Makantasis, Georgios N. Yannakakis

机构 * Institute of Digital Games University of Malta(数字游戏研究所马耳他大学) Department of Artificial Intelligence University of Malta(人工智能系马耳他大学)

专题命中 跨模态检索 :multimodal(title);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23736 2025-08-07 cs.CV cs.IR 57%

Modality and Task Adaptation for Enhanced Zero-shot Composed Image Retrieval

Haiwen Li, Fei Su, Zhicheng Zhao

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 11 篇

2412.03859 2025-08-07 cs.CV 83%

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Hui Zhang, Dexiang Hong, Yitong Wang, Jie Shao, Xinglong Wu, Zuxuan Wu, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI(可信具身人工智能研究院) Fudan University(复旦大学) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) Bytedance Intelligent Creation(字节跳动智能创作)

专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21167 2025-08-07 cs.CV cs.AI 81%

ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions

Donglu Yang, Liang Zhang, Zihao Yue, Liangyu Chen, Yichen Xu, Wenxuan Wang, Qin Jin

机构 * independent researcher(独立研究者)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05683 2025-08-07 cs.LG cs.AI cs.CR cs.MM 81%

Multi-Modal Multi-Task Federated Foundation Models for Next-Generation Extended Reality Systems: Towards Privacy-Preserving Distributed Intelligence in AR/VR/MR

Fardis Nadimi, Payam Abdisarabshali, Kasra Borazjani, Jacob Chakareski, Seyyedali Hosseinalipour

机构 * University at Buffalo–SUNY(布法罗大学-纽约州立大学) Department of Electrical Engineering(电气工程系) New Jersey Institute of Technology (NJIT)(新泽西理工学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI、cs.MM

Comments 16 pages, 4 Figures, 8 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04229 2025-08-07 cs.CV 79%

Intention Enhanced Diffusion Model for Multimodal Pedestrian Trajectory Prediction

Yu Liu, Zhijie Liu, Xiao Ren, You-Fu Li, He Kong

机构 * Guangdong Provincial Key Laboratory of Fully Actuated System Control Theory and Technology, the Southern University of Science and Technology, Shenzhen 518055, China(广东省全自动化系统控制理论与技术重点实验室,南方科技大学,深圳518055,中国) Department of Mechanical Engineering, City University of Hong Kong, Hong Kong SAR, China(香港城市大学机械工程系,香港特别行政区,中国)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments To be presented at the 28th IEEE International Conference on Intelligent Transportation Systems (ITSC), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20830 2025-08-07 cs.CV 79%

CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation

Jianyu Wu, Yizhou Wang, Xiangyu Yue, Xinzhu Ma, Jingyang Guo, Dongzhan Zhou, Wanli Ouyang, Shixiang Tang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Beihang University(北京航空航天大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04271 2025-08-07 cs.DC 78%

S2M3: Split-and-Share Multi-Modal Models for Distributed Multi-Task Inference on the Edge

JinYi Yoon, JiHo Lee, Ting He, Nakjung Choi, Bo Ji

专题命中 多模态生成 :multi-modal(title,abstract)

Comments Accepted at IEEE International Conference on Distributed Computing Systems (ICDCS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18087 2025-08-07 cs.CV 70%

Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation

Weipeng Tan, Chuming Lin, Chengming Xu, FeiFan Xu, Xiaobin Hu, Xiaozhong Ji, Junwei Zhu, Chengjie Wang, Yanwei Fu

机构 * Fudan University(复旦大学) Tencent, YouTu Lab(腾讯、YouTu实验室)

专题命中 多模态生成 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments Accepted by ACM MM'25. arXiv admin note: text overlap with arXiv:2409.03270

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03696 2025-08-07 cs.CR cs.AI cs.CV 62%

PLA: Prompt Learning Attack against Text-to-Image Generative Models

Xinqi Lyu, Yihao Liu, Yanjie Li, Bin Xiao

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 3 figures, and published to ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏