arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-08 至 2025-08-08 共收录 53 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4 篇

2508.05167 2025-08-08 cs.CV 83%

PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-based Autonomous Driving Systems

Qi Guo, Xiaojun Jia, Shanmin Pang, Simeng Qin, Lin Wang, Ju Jia, Yang Liu, Qing Guo

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05383 2025-08-08 cs.AI 79%

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models

Xiangxiang Zhang, Jingxuan Wei, Donghong Zhong, Qi Chen, Caijun Jia, Cheng Tan, Jinming Gu, Xiaobo Qin, Zhiping Liu, Liang Hu, Tong Sun, Yuchen Wu, Zewei Sun, Chenwei Lou, Hua Zheng, Tianyang Zhan, Changbao Wang, Shuangzhi Wu, Zefa Lin, Chang Guo, Sihang Yuan, Riwei Chen, Shixiong Zhao, Yingping Zhang, Gaowei Wu, Bihui Yu, Jiahui Wu, Zhehui Zhao, Qianqian Liu, Ruofeng Tang, Xingyue Huang, Bing Zhao, Mengyang Zhang, Youqiang Zhou

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04868 2025-08-08 cs.CV 79%

Dual-Stream Attention with Multi-Modal Queries for Object Detection in Transportation Applications

Noreen Anwar, Guillaume-Alexandre Bilodeau, Wassim Bouachir

机构 * LITIV, Polytechnique Montréal(Polytechnique Montréal 的 LITIV) Data Science Laboratory, Université du Québec (TELUQ)(Université du Québec (TELUQ) 的 Data Science Laboratory)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04987 2025-08-08 cs.CV 57%

Unified modality separation: A vision-language framework for unsupervised domain adaptation

Xinyao Li, Jingjing Li, Zhekai Du, Lei Zhu, Heng Tao Shen

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) Tongji University(同济大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 5 篇

2508.00733 2025-08-08 cs.SD cs.CV cs.MM eess.AS 85%

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Le Wang, Jun Wang, Chunyu Qiang, Feng Deng, Chen Zhang, Di Zhang, Kun Gai

机构 * China University of Mining and Technology(中国矿业大学) Kuaishou Technology(快手科技)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments 12 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05087 2025-08-08 cs.MM cs.AI cs.CL cs.CR 82%

JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering

Renmiao Chen, Shiyao Cui, Xuancheng Huang, Chengwei Pan, Victor Shea-Jay Huang, QingLin Zhang, Xuan Ouyang, Zhexin Zhang, Hongning Wang, Minlie Huang

机构 * CoAI group, DCST, Tsinghua University(清华大学DCST学院) Beihang University(北航大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments 10 pages, 3 tables, 2 figures, to appear in the Proceedings of the 33rd ACM International Conference on Multimedia (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04723 2025-08-08 cs.SD cs.AI eess.AS 73%

Wearable Music2Emotion : Assessing Emotions Induced by AI-Generated Music through Portable EEG-fNIRS Fusion

Sha Zhao, Song Yi, Yangxuan Zhou, Jiadong Pan, Jiquan Wang, Jie Xia, Shijian Li, Shurong Dong, Gang Pan

机构 * Zhejiang University(浙江大学) Hangzhou RongNao Technology Co., Ltd(杭州融脑科技有限公司)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI、eess.AS

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04585 2025-08-08 eess.AS 57%

UniTalker: Conversational Speech-Visual Synthesis

Yifan Hu, Rui Liu, Yi Ren, Xiang Yin, Haizhou Li

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 15 pages, 8 figures, Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03536 2025-08-08 eess.AS 57%

Overview of Automatic Speech Analysis and Technologies for Neurodegenerative Disorders: Diagnosis and Assistive Applications

Shakeel A. Sheikh, Md. Sahidullah, Ina Kodrasi

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Published in IEEE Journal of Selected Topics in Signal Processing

Journal ref https://ieeexplore.ieee.org/abstract/document/11086511/

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2411.08882 2025-08-08 cs.MM cs.AI cs.CV 82%

A Novel Multimodal System to Predict Agitation in People with Dementia Within Clinical Settings: A Proof of Concept

Abeer Badawi, Somayya Elmoghazy, Samira Choudhury, Sara Elgazzar, Khalid Elgazzar, Amer Burhan

机构 * IoT Research Laboratory, Ontario Tech University(Ontario Tech 大学物联网研究实验室) Ontario Shores Centre for Mental Health Sciences(Ontario Shores 精神健康科学中心) Temerty Faculty of Medicine, University of Toronto(多伦多大学Temerty医学学院) Faculty of Science, Ontario Tech University(Ontario Tech 大学科学学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04900 2025-08-08 cs.CV cs.AI 81%

Revealing Temporal Label Noise in Multimodal Hateful Video Classification

Shuonan Yang, Tailin Chen, Rahul Singh, Jiangbei Yue, Jianbo Jiao, Zeyu Fu

机构 * Multimodal Intelligence Lab(多模态智能实验室) Department of Computer Science(计算机科学系) University of Exeter(埃克塞特大学) University of Birmingham(伯明翰大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21586 2025-08-08 cs.CL cs.AI cs.CV 67%

Can Vision Language Models Understand Mimed Actions?

Hyundong Cho, Spencer Lin, Tejas Srinivasan, Michael Saxon, Deuksin Kwon, Natali T. Chavez, Jonathan May

机构 * Information Sciences Institute(信息科学研究所) Institute for Creative Technologies(创意技术研究所) Department of Computer Science(计算机科学系) University of Southern California(南加州大学) University of California, Santa Barbara(加州大学圣巴巴拉分校) Aristotle University of Thessaloniki(希腊雅典纳大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05145 2025-08-08 cs.AI 57%

Graph-based Event Log Repair

Sebastiano Dissegna, Chiara Di Francescomarino, Massimiliano Ronzani

机构 * Department of Computer Science Engineering University of Trento(计算机科学工程系 特伦托大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10855 2025-08-08 cs.RO cs.LG 50%

Fast and Robust Visuomotor Riemannian Flow Matching Policy

Haoran Ding, Noémie Jaquier, Jan Peters, Leonel Rozo

机构 * Bosch Center for Artificial Intelligence(博世人工智能中心) Division of Robotics, Perception, and Learning, KTH Royal Institute of Technology(机器人、感知与学习 division,皇家理工学院) Computer Science Department of the Technische Universität Darmstadt(达姆施塔特技术大学计算机科学系)

专题命中 视频多模态 :multi-modal(abstract)

Comments Accepted for publication in IEEE T-RO. Project website: https://sites.google.com/view/rfmp 17 pages, 12 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2502.17297 2025-08-08 cs.AI 79%

Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts

Zhenghao Liu, Xingsheng Zhu, Tianshuo Zhou, Xinyi Zhang, Xiaoyuan Yi, Yukun Yan, Ge Yu, Maosong Sun

机构 * Northeastern University, China(东北大学) Microsoft Research Asia(微软亚洲研究院) Tsinghua University(清华大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23029 2025-08-08 cs.CL 57%

Uncovering Visual-Semantic Psycholinguistic Properties from the Distributional Structure of Text Embedding Space

Si Wu, Sebastian Bruch

机构 * Northeastern University(东北大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

Comments The camera-ready version for ACL 2025 in Vienna

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 8 篇

2508.05234 2025-08-08 cs.CL cs.AI 84%

Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation

Haonan Shangguan, Xiaocui Yang, Shi Feng, Daling Wang, Yifei Zhang, Ge Yu

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05352 2025-08-08 cs.IR cs.AI 83%

Multi-Modal Multi-Behavior Sequential Recommendation with Conditional Diffusion-Based Feature Denoising

Xiaoxi Cui, Weihai Lu, Yu Tong, Yiheng Li, Zhejun Zhao

机构 * Peking University(北京大学) Wuhan University(武汉大学) Shanghai University of International Business(上海国际商务大学) Microsoft(微软公司)

专题命中 多模态生成 :multi-modal(title,abstract);multimodal(abstract);分类 cs.AI

Comments SIGIR 2025

Journal ref SIGIR 2025: Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval Pages 1593 - 1602

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17963 2025-08-08 cs.CV 83%

M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation

Xiaowei Chi, Junbo Qi, Rongyu Zhang, Shanghang Zhang, Qifeng Liu, Yike Guo

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Waseda University(早稻田大学) Peking University(北京大学)

专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06510 2025-08-08 cs.CV cs.AI 81%

AnomalyControl: Learning Cross-modal Semantic Features for Controllable Anomaly Synthesis

Shidan He, Lei Liu, Xiujun Shu, Bo Wang, Yuanhao Feng, Shen Zhao

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03069 2025-08-08 cs.CV cs.AI 81%

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Liao Qu, Huichao Zhang, Yiheng Liu, Xu Wang, Yi Jiang, Yiming Gao, Hu Ye, Daniel K. Du, Zehuan Yuan, Xinglong Wu

机构 * ByteDance(字节跳动)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments CVPR 2025; Code and models: https://github.com/ByteVisionLab/TokenFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02317 2025-08-08 cs.CL cs.AI cs.DC 62%

VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo

Qianli Ma, Yaowei Zheng, Zhelun Shi, Zhongkai Zhao, Bin Jia, Ziyue Huang, Zhiqi Lin, Youjie Li, Jiacheng Yang, Yanghua Peng, Zhi Zhang, Xin Liu

机构 * ByteDance(字节跳动)

专题命中 多模态生成 :omni-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16978 2025-08-08 cs.CV cs.AI 62%

PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask

Jeongho Kim, Hoiyeong Jin, Sunghyun Park, Jaegul Choo

机构 * KAIST(韩国科学技术院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04732 2025-08-08 cs.LG cs.GR 50%

LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation

Xiaoqi Dong, Xiangyu Zhou, Nicholas Evans, Yujia Lin

机构 * Dali University(大理大学) Bandırma Onyedi Eylül University(班迪尔马第十七个九月大学)

专题命中 多模态生成 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 13 篇

2508.05602 2025-08-08 cs.CV 90%

LLaVA-RE: Binary Image-Text Relevancy Evaluation with Multimodal Large Language Model

Tao Sun, Oliver Liu, JinJin Li, Lan Ma

机构 * Stony Brook University(斯通布罗克大学) Amazon(亚马逊)

专题命中 多模态评测 :multimodal(title,abstract);image-text(title,abstract);MLLM(abstract);分类 cs.CV

Comments Published in the First Workshop of Evaluation of Multi-Modal Generation 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07689 2025-08-08 cs.CV cs.MM cs.RO 81%

RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving

Zhijian Huang, Chengjian Feng, Feng Yan, Baihui Xiao, Zequn Jie, Yujie Zhong, Xiaodan Liang, Lin Ma

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Meituan(美团)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05527 2025-08-08 cs.CV 79%

AI vs. Human Moderators: A Comparative Evaluation of Multimodal LLMs in Content Moderation for Brand Safety

Adi Levi, Or Levi, Sardhendu Mishra, Jonathan Morra

机构 * Zefr Inc(Zefr公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to the Computer Vision in Advertising and Marketing (CVAM) workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05083 2025-08-08 cs.AI 79%

MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models

Dexuan Xu, Jieyi Wang, Zhongyan Chai, Yongzhi Cao, Hanpin Wang, Huamin Zhang, Yu Huang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05053 2025-08-08 cs.CV 79%

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?

Parth Thakkar, Ankush Agarwal, Prasad Kasu, Pulkit Bansal, Chaitanya Devaguptapu

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Accepted at ACL 2025 in the main track

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04088 2025-08-08 cs.CL 79%

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning

Jianghangfan Zhang, Yibo Yan, Kening Zheng, Xin Zou, Song Dai, Xuming Hu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏