arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-15 至 2025-08-15 共收录 59 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 7 篇

2508.10110 2025-08-15 cs.CV cs.AI 81%

Empowering Morphing Attack Detection using Interpretable Image-Text Foundation Model

Sushrut Patwardhan, Raghavendra Ramachandra, Sushma Venkatesh

机构 * Norwegian University of Science and Technology (NTNU)(挪威科学技术大学) MOBAI AS(MOBAI公司)

专题命中 图文多模态 :image-text(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10444 2025-08-15 cs.CL 79%

DiFaR: Enhancing Multimodal Misinformation Detection with Diverse, Factual, and Relevant Rationales

Herun Wan, Jiaying Wu, Minnan Luo, Xiangzheng Kong, Zihan Ma, Zhi Zeng

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14976 2025-08-15 cs.CV 79%

Hierarchical Cross-modal Prompt Learning for Vision-Language Models

Hao Zheng, Shunzhi Yang, Zhuoxin He, Jinfeng Yang, Zhenhua Huang

机构 * South China Normal University(华南师范大学) Shenzhen Polytechnic University(深圳职业技术大学)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10339 2025-08-15 cs.CV cs.LG 74%

Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models

Andrew Bai, Justin Cui, Ruochen Wang, Cho-Jui Hsieh

机构 * Department of Computer Science University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

Comments 11 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10645 2025-08-15 cs.CV 57%

SemPT: Semantic Prompt Tuning for Vision-Language Models

Xiao Shi, Yangjun Ou, Zhenzhong Chen

机构 * School of Computer Science and Artificial Intelligence, Wuhan Textile University(武汉纺织大学计算机科学与人工智能学院) School of Remote Sensing and Information Engineering, Wuhan University(武汉大学遥感与信息工程学院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10333 2025-08-15 cs.RO cs.CV 57%

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

Wenxuan Song, Ziyang Zhou, Han Zhao, Jiayi Chen, Pengxiang Ding, Haodong Yan, Yuxin Huang, Feilong Tang, Donglin Wang, Haoang Li

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07985 2025-08-15 cs.CV 57%

Common Data Properties Limit Object-Attribute Binding in CLIP

Bijay Gurung, David T. Hoffmann, Thomas Brox

机构 * University of Freiburg(弗赖堡大学) deepset

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments accepted at GCPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 5 篇

2506.23009 2025-08-15 cs.CV 83%

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

Jian Chen, Wenye Ma, Penghang Liu, Wei Wang, Tengwei Song, Ming Li, Chenguang Wang, Jiayu Qin, Ruiyi Zhang, Changyou Chen

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06117 2025-08-15 cs.CL cs.AI 62%

Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding

Fabian David Schmidt, Ivan Vulić, Goran Glavaš, David Ifeoluwa Adelani

机构 * Center For Artificial Intelligence and Data Science, University of Würzburg(人工智能与数据科学中心,乌尔姆大学) Language Technology Lab, University of Cambridge(语言技术实验室,剑桥大学) Mila, McGill University and Canada CIFAR AI Chair(Mila,麦吉尔大学及加拿大CIFAR人工智能主席)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10830 2025-08-15 cs.SD eess.AS 57%

Advances in Speech Separation: Techniques, Challenges, and Future Trends

Kai Li, Guo Chen, Wendi Sang, Yi Luo, Zhuo Chen, Shuai Wang, Shulin He, Zhong-Qiu Wang, Andong Li, Zhiyong Wu, Xiaolin Hu

机构 * Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学) School of Computer Technology and Application, Qinghai University(计算机技术与应用学院,青海大学) ByteDance(字节跳动) Nanjing University(南京大学) Southern University of Science and Technology(南方科技大学) Institute of Acoustics, Chinese Academy of Sciences(中国科学院声学研究所)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 34 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10580 2025-08-15 cs.MM cs.SD 57%

Ensembling Synchronisation-based and Face-Voice Association Paradigms for Robust Active Speaker Detection in Egocentric Recordings

Jason Clarke, Yoshihiko Gotoh, Stefan Goetze

机构 * Speech and Hearing (SPandH), School of Computer Science, The University of Sheffield, UK(语音与听力(SPandH)、计算机科学学院、谢菲尔德大学、英国) South Westphalia University of Applied Sciences, Iserlohn, Germany(西南弗兰肯应用科学大学、伊塞尔洛恩、德国)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM

Comments Accepted to SPECOM 2025, 13 pages, 4 figures. To appear in the Proceedings of the 27th International Conference on Speech and Computer (SPECOM) 2025, October 13-14, 2025, Szeged, Hungary

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10130 2025-08-15 q-bio.NC 50%

Linking GFAP Levels to Speech Anomalies in Acute Brain Injury: A Simulation Based Study

Shamaley Aravinthan, Bin Hu

专题命中 音频语音多模态 :multimodal(abstract)

Comments 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 6 篇

2508.10572 2025-08-15 cs.CV 83%

Towards Agentic AI for Multimodal-Guided Video Object Segmentation

Tuyen Tran, Thao Minh Le, Truyen Tran

机构 * Applied Artificial Intelligence Institute, Deakin University, Australia(应用人工智能研究所,德金大学,澳大利亚)

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10828 2025-08-15 cs.RO cs.AI 79%

A Multimodal Neural Network for Recognizing Subjective Self-Disclosure Towards Social Robots

Henry Powell, Guy Laban, Emily S. Cross

机构 * Amazon(亚马逊) School of Psychology and Neuroscience, University of Glasgow(心理学与神经科学学院,格拉斯哥大学) Ben-Gurion University of the Negev(内盖夫本·古里安大学) University of Cambridge(剑桥大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10784 2025-08-15 q-bio.NC cs.CV 57%

Insights from the Algonauts 2025 Winners

Paul S. Scotti, Mihir Tripathy

机构 * Medical AI Research Center ( MedARC )(医学人工智能研究中心(MedARC)) Baylor College of Medicine(贝勒医学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Perspective piece on Algonauts 2025 Challenge conclusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07171 2025-08-15 cs.CV 57%

EventRR: Event Referential Reasoning for Referring Video Object Segmentation

Huihui Xu, Jiashi Lin, Haoyu Chen, Junjun He, Lei Zhu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) Northwestern Polytechnical University(西北工业大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00399 2025-08-15 cs.CV 57%

iSafetyBench: A video-language benchmark for safety in industrial environment

Raiyaan Abdullah, Yogesh Singh Rawat, Shruti Vyas

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to VISION'25 - ICCV 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03096 2025-08-15 cs.CV 57%

Scaling Open-Vocabulary Action Detection

Zhen Hao Sia, Yogesh Singh Rawat

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 4 篇

2507.14189 2025-08-15 cs.CL cs.AI 81%

DeepWriter: A Fact-Grounded Multimodal Writing Assistant Based On Offline Knowledge Base

Song Mao, Lejun Cheng, Pinlong Cai, Guohang Yan, Ding Wang, Botian Shi

机构 * Shanghai AI LAB(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments work in process

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06071 2025-08-15 cs.CV cs.MM 81%

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding

Chang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja, Xiaokang Yang

机构 * Shanghai Jiao Tong University(上海交通大学) Singapore Institute of Technology(新加坡科技学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14493 2025-08-15 cs.IR cs.AI cs.LG 57%

FinSage: A Multi-aspect RAG System for Financial Filings Question Answering

Xinyu Wang, Jijun Chi, Zhenghan Tai, Tung Sum Thomas Kwok, Muzhi Li, Zhuhong Li, Hailin He, Yuchen Hua, Peng Lu, Suyuchen Wang, Yihong Wu, Jerry Huang, Jingrui Tian, Fengran Mo, Yufei Cui, Ling Zhou

机构 * 1SimpleWay.AI 2McGill University 3University of Toronto 4University of California, Los Angeles 5The Chinese University of Hong Kong 6Duke University 7Universit\'e de Montr\'eal 8Mila - Quebec AI Institute 9Noah's Ark Lab 10CG Matrix Technology Limited 1SimpleWay.AI 2McGill University 3University of Toronto 4University of California, Los Angeles 5The Chinese University of Hong Kong 6Duke University 7Universit\'e de Montr\'eal 8Mila - Quebec AI Institute 9Noah's Ark Lab 10CG Matrix Technology Limited

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

Comments Accepted at the 34th ACM International Conference on Information and Knowledge Management (CIKM2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04376 2025-08-15 cs.CV 57%

MIDAS: Modeling Ground-Truth Distributions with Dark Knowledge for Domain Generalized Stereo Matching

Peng Xu, Zhiyu Xiang, Jingyun Fu, Tianyu Pu, Hanzhi Zhong, Eryun Liu

机构 * Zhejiang University, China(浙江大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 10 篇

2508.10494 2025-08-15 cs.LG cs.AI cs.MA 85%

A Unified Multi-Agent Framework for Universal Multimodal Understanding and Generation

Jiulin Li, Ping Huang, Yexin Li, Shuo Chen, Juewen Hu, Ye Tian

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);any-to-any(abstract);分类 cs.AI

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05262 2025-08-15 cs.CV 79%

Debiasing Multimodal Large Language Models via Penalization of Language Priors

YiFan Zhang, Yang Shi, Weichen Yu, Qingsong Wen, Xue Wang, Wenjing Yang, Zhang Zhang, Liang Wang, Rong Jin

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Peking University(北京大学) Carnegie Mellon University(卡内基梅隆大学) Alibaba Group(阿里巴巴集团) National University of Defense Technology(国防科技大学) Meta

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08098 2025-08-15 cs.CV 70%

TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning

Junzhe Xu, Yuyang Yin, Xi Chen

机构 * Basic Algorithm Center, PCG, Tencent(腾讯基础算法中心、PCG、腾讯)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13307 2025-08-15 cs.CV cs.AI 62%

Quantitative Comparison of Fine-Tuning Techniques for Pretrained Latent Diffusion Models in the Generation of Unseen SAR Images

Solène Debuysère, Nicolas Trouvé, Nathan Letheule, Olivier Lévêque, Elise Colin

机构 * Paris-Saclay University(巴黎-萨克雷大学) ONERA - The French Aerospace Lab(法国航空航天实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08601 2025-08-15 cs.CV cs.AI 62%

Yan: Foundational Interactive Video Generation

Deheng Ye, Fangyun Zhou, Jiacheng Lv, Jianqi Ma, Jun Zhang, Junyan Lv, Junyou Li, Minwen Deng, Mingyu Yang, Qiang Fu, Wei Yang, Wenkai Lv, Yangbin Yu, Yewen Wang, Yonghang Guan, Zhihao Hu, Zhongbin Fang, Zhongqian Sun

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10858 2025-08-15 cs.CV 57%

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation

Harold Haodong Chen, Haojian Huang, Qifeng Chen, Harry Yang, Ser-Nam Lim

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Project Page: https://haroldchen19.github.io/PhysHPO-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10280 2025-08-15 cs.CV 57%

High Fidelity Text to Image Generation with Contrastive Alignment and Structural Guidance

Danyi Gao

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10013 2025-08-15 cs.CL 57%

Semantic Bridge: Universal Multi-Hop Question Generation via AMR-Driven Graph Synthesis

Linqing Chen, Hanmeng Zhong, Wentao Wu, Weilei Wang

机构 * Linqing Chen, Hanmeng Zhong, Wentao Wu, Weilei Wang(作者)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏