arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-16 至 2025-09-16 共收录 74 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 10 篇

2501.15688 2025-09-16 cs.CL cs.AI cs.LG 84%

Transformer-Based Multimodal Knowledge Graph Completion with Link-Aware Contexts

Haodi Ma, Dzmitry Kasinets, Daisy Zhe Wang

机构 * Department of Computer and Information Science and Engineering, University of Florida(计算机与信息科学与工程系,佛罗里达大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00284 2025-09-16 cs.RO cs.AI 79%

LightEMMA: Lightweight End-to-End Multimodal Model for Autonomous Driving

Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu

机构 * Department of Civil and Environmental Engineering, University of Michigan(土木与环境工程系,密歇根大学) University of Michigan Transportation Research Institute(密歇根大学交通研究所)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11247 2025-09-16 cs.CV 74%

Contextualized Multimodal Lifelong Person Re-Identification in Hybrid Clothing States

Robert Long, Rongxin Jiang, Mingrui Yan

机构 * University of Padua(帕多瓦大学) Heilongjiang University of Science and Technology(黑龙江科技大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01064 2025-09-16 cs.CV cs.AI 73%

Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs

Yudong Zhang, Ruobing Xie, Yiqing Huang, Jiansheng Chen, Xingwu Sun, Zhanhui Kang, Di Wang, Yu Wang

机构 * Tsinghua University, Tencent(清华大学,腾讯) Tencent(腾讯) University of Science and Technology Beijing(北京科技大学) Tencent, University of Macau(腾讯,澳门大学) Tsinghua University(清华大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM Multimedia 2025 BNI track (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16915 2025-09-16 cs.CV cs.LG 70%

Multilingual Diversity Improves Vision-Language Representations

Thao Nguyen, Matthew Wallingford, Sebastin Santy, Wei-Chiu Ma, Sewoong Oh, Ludwig Schmidt, Pang Wei Koh, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Allen Institute for Artificial Intelligence(人工智能研究院)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments NeurIPS 2024 Spotlight paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16146 2025-09-16 cs.CV cs.AI cs.CL cs.LG 67%

Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation

Zhenglin Hua, Jinghan He, Zijun Yao, Tianxu Han, Haiyun Guo, Yuheng Jia, Junfeng Fang

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University)(东南大学新一代人工智能技术及其交叉应用关键实验室) Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Wuhan University of Technology(武汉理工大学) National University of Singapore(新加坡国立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11895 2025-09-16 cs.CV cs.AI 62%

Integrating Prior Observations for Incremental 3D Scene Graph Prediction

Marian Renz, Felix Igelbrink, Martin Atzmueller

机构 * DFKI Niedersachsen(德克萨斯联合研究所(北莱茵威斯特法伦)) Cooperative and Autonomous Systems, DFKI Niedersachsen(合作与自主系统,DFKI北莱茵威斯特法伦) German Research Center for Artificial Intelligence(德国人工智能研究中心) Semantic Information Systems, Osnabrück University(语义信息系统,奥斯纳布吕克大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at 24th International Conference on Machine Learning and Applications (ICMLA'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11961 2025-09-16 cs.CL 57%

Spec-LLaVA: Accelerating Vision-Language Models with Dynamic Tree-Based Speculative Decoding

Mingxiao Huo, Jiayi Zhang, Hewei Wang, Jinfeng Xu, Zheyu Chen, Huilin Tai, Yijun Chen

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Nottingham(诺丁汉大学) The University of Hong Kong(香港大学) The Hong Kong Polytechnic University(香港理工大学) Columbia University(哥伦比亚大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

Comments 7pages, accepted by ICML TTODLer-FM workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05479 2025-09-16 cs.CV 57%

LATTE: Learning to Think with Vision Specialists

Zixian Ma, Jianguo Zhang, Zhiwei Liu, Jieyu Zhang, Juntao Tan, Manli Shu, Juan Carlos Niebles, Shelby Heinecke, Huan Wang, Caiming Xiong, Ranjay Krishna, Silvio Savarese

机构 * University of Washington(华盛顿大学) Salesforce Research(Salesforce研究)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11065 2025-09-16 cs.SE cs.PL 50%

ViScratch: Using Large Language Models and Gameplay Videos for Automated Feedback in Scratch

Yuan Si, Daming Li, Hanyuan Shi, Jialu Zhang

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 6 篇

2509.11937 2025-09-16 cs.SE cs.AI 79%

MMORE: Massive Multimodal Open RAG & Extraction

Alexandre Sallinen, Stefan Krsteski, Paul Teiletche, Marc-Antoine Allard, Baptiste Lecoeur, Michael Zhang, Fabrice Nemo, David Kalajdzic, Matthias Meyer, Mary-Anne Hartley

机构 * École Polytechnique Fédérale de Lausanne (EPFL), Switzerland(瑞士联邦理工学院洛桑校区) ETH Zürich, Switzerland(瑞士苏黎世联邦理工学院) T.H. Chan School of Public Health, Harvard University, USA(哈佛大学T.H. Chan公共卫生学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments This paper was originally submitted to the CODEML workshop for ICML 2025. 9 pages (including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11183 2025-09-16 cs.SD eess.AS 79%

WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation

Emmanouil Karystinaios

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Accepted at Large Language Models for Music & Audio Workshop (LLM4MA) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18039 2025-09-16 cs.AI 79%

MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind

Zheng Zhang, Nuoqian Xiao, Qi Chai, Deheng Ye, Hao Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tencent(腾讯)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12204 2025-09-16 cs.CV 70%

Character-Centric Understanding of Animated Movies

Zhongrui Gui, Junyu Xie, Tengda Han, Weidi Xie, Andrew Zisserman

机构 * Visual Geometry Group Dept.\ of Engineering Science University of Oxford, UK School of Artifitial Intelligence\ Jiao Tong University Shanghai China School of Artifitial Intelligence\ Jiao Tong University

专题命中 音频语音多模态 :multi-modal(abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08638 2025-09-16 eess.AS cs.AI cs.MM cs.SD 56%

YuE: Scaling Open Foundation Models for Long-Form Music Generation

Ruibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang, Jiahao Pan, Yongyi Zang, Haohe Liu, Yiming Liang, Wenye Ma, Xingjian Du, Xinrun Du, Zhen Ye, Tianyu Zheng, Zhengxuan Jiang, Yinghao Ma, Minghao Liu, Zeyue Tian, Ziya Zhou, Liumeng Xue, Xingwei Qu, Yizhi Li, Shangda Wu, Tianhao Shen, Ziyang Ma, Jun Zhan, Chunhui Wang, Yatian Wang, Xiaowei Chi, Xinyue Zhang, Zhenzhu Yang, Xiangzhou Wang, Shansong Liu, Lingrui Mei, Peng Li, Junjie Wang, Jianwei Yu, Guojian Pang, Xu Li, Zihao Wang, Xiaohuan Zhou, Lijun Yu, Emmanouil Benetos, Yong Chen, Chenghua Lin, Xie Chen, Gus Xia, Zhaoxiang Zhang, Chao Zhang, Wenhu Chen, Xinyu Zhou, Xipeng Qiu, Roger Dannenberg, Jiaheng Liu, Jian Yang, Wenhao Huang, Wei Xue, Xu Tan, Yike Guo

机构 * HKUST(香港科技大学) MAP(多模态艺术投影)

专题命中 音频语音多模态 :分类 cs.AI、cs.MM、eess.AS;multimodal(comments)

Comments https://github.com/multimodal-art-projection/YuE

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02051 2025-09-16 cs.CR 50%

Sanitization of Multimedia Content: A Survey of Techniques, Attacks, and Future Directions

Andrea Ciccotelli, Hanaa Abbas, Roberto Di Pietro

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 8 篇

2406.13763 2025-09-16 cs.CV cs.AI 81%

Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models

Zhawnen Chen, Tianchun Wang, Yizhou Wang, Michal Kosinski, Xiang Zhang, Yun Fu, Sheng Li

机构 * University of Virginia(弗吉尼亚大学) The Pennsylvania State University(宾夕法尼亚州立大学) Northeastern University(东北大学) Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11944 2025-09-16 cs.AI 79%

Agentic Temporal Graph of Reasoning with Multimodal Language Models: A Potential AI Aid to Healthcare

Susanta Mitra

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10802 2025-09-16 q-fin.RM cs.CL cs.LG q-fin.CP 79%

Why Bonds Fail Differently? Explainable Multimodal Learning for Multi-Class Default Prediction

Yi Lu, Aifan Ling, Chaoqun Wang, Yaxin Xu

机构 * School of Economics and Finance, Shanghai International Studies University(经济金融学院,上海国际问题研究大学) School of AI and Advanced Computing, Xi’an Jiaotong-Liverpool University(人工智能与先进计算学院,西安交通大学利物浦大学) School of Foreign Studies, Shanghai University of Finance and Economics(外国语言学院,上海金融学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11232 2025-09-16 cs.CV cs.AI 62%

MIS-LSTM: Multichannel Image-Sequence LSTM for Sleep Quality and Stress Prediction

Seongwan Park, Jieun Woo, Siheon Yang

机构 * Sungkyunkwan University(釜山大学) Yeungnam University(延世大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ICTC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11796 2025-09-16 cs.CV 57%

FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning

Haodong Chen, Haojian Huang, XinXiang Yin, Dian Shao

机构 * School of Automation, Northwestern Polytechnical University(自动化学院,西北工业大学) The University of Hong Kong(香港大学) School of Software, Northwestern Polytechnical University(软件学院,西北工业大学) Unmanned System Research Institute, Northwestern Polytechnical University(无人系统研究院,西北工业大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11713 2025-09-16 cs.LG cs.NI 50%

Beyond Regularity: Modeling Chaotic Mobility Patterns for Next Location Prediction

Yuqian Wu, Yuhong Peng, Jiapeng Yu, Xiangyu Liu, Zeting Yan, Kang Lin, Weifeng Su, Bingqing Qu, Raymond Lee, Dingqi Yang

机构 * Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist大学) University of Warwick(沃里克大学) University of Macau(澳门大学)

专题命中 视频多模态 :multimodal(abstract)

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10864 2025-09-16 cs.LG 50%

CogGNN: Cognitive Graph Neural Networks in Generative Connectomics

Mayssa Soussia, Yijun Lin, Mohamed Ali Mahjoub, Islem Rekik

机构 * National Engineering School of Sousse, University of Sousse, LATIS- Laboratory of Advanced Technology and Intelligent Systems(突尼斯苏塞国立工程学校,苏塞大学,先进技术与智能系统实验室) BASIRA Lab, Imperial-X(BASIRA实验室,Imperial-X) Department of Computing, Imperial College London, UK(计算系,伦敦帝国学院,英国)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10552 2025-09-16 q-bio.NC cs.LG 50%

Trial-Level Time-frequency EEG Desynchronization as a Neural Marker of Pain

D. A. Blanco-Mora, A. Dierolf, J. Gonçalves, M. van Der Meulen

机构 * Luxembourg Centre for Systems Biomedicine, University of Luxembourg(卢森堡系统生物医学研究中心,卢森堡大学) Department of Behavioural and Cognitive Sciences, University of Luxembourg(行为与认知科学系,卢森堡大学)

专题命中 视频多模态 :multimodal(abstract)

Comments 7 pages, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 4 篇

2509.11054 2025-09-16 cs.IT cs.CV math.IT 83%

Rate-Distortion Limits for Multimodal Retrieval: Theory, Optimal Codes, and Finite-Sample Guarantees

Thomas Y. Chen

机构 * Department of Computer Science, Columbia University(计算机科学系,哥伦比亚大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments ICCV MRR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10467 2025-09-16 cs.IR cs.AI cs.CL cs.CV cs.MM 83%

DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph

Mengzheng Yang, Yanfei Ren, David Osei Opoku, Ruochang Li, Peng Ren, Chunxiao Xing

机构 * School of Software, Henan University, Kaifeng 475004, China(河南大学软件学院) BNRist, DCST, RIIT, Tsinghua University, Beijing 100084, China(清华大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 12 pages, 5 figures. Accepted to the 22nd International Conference on Web Information Systems and Applications (WISA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11587 2025-09-16 cs.CV cs.AI 62%

Hierarchical Identity Learning for Unsupervised Visible-Infrared Person Re-Identification

Haonan Shi, Yubin Wang, De Cheng, Lingfeng He, Nannan Wang, Xinbo Gao

机构 * IEEE Publication Technology Department(IEEE出版技术部门) State Key Laboratory of Integrated Services Networks, School of Telecommunications Engineering, Xidian University(信息服务网络国家重点实验室,电信工程学院,西安电子科技大学) Department of Computer Science and Technology, Tongji University(计算机科学与技术系,同济大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10697 2025-09-16 cs.CL 57%

A Survey on Retrieval And Structuring Augmented Generation with Large Language Models

Pengcheng Jiang, Siru Ouyang, Yizhu Jiao, Ming Zhong, Runchu Tian, Jiawei Han

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments KDD'25 survey track

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 13 篇

2406.10424 2025-09-16 cs.CV cs.AI 84%

What is the Visual Cognition Gap between Humans and Multimodal LLMs?

Xu Cao, Yifan Shen, Bolin Lai, Wenqian Ye, Yunsheng Ma, Joerg Heintz, Jintai Chen, Meihuan Huang, Jianguo Cao, Aidong Zhang, James M. Rehg

机构 * Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系) College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院) Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系) Digital Twin Lab, Purdue University(普渡大学数字孪生实验室) HKUST (Guangzhou)(香港科技大学(广州)) Department of Rehabilitation Medicine, Shenzhen Children’s Hospital(深圳儿童医院康复医学系)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09070 2025-09-16 cs.LG cs.AI cs.CV 81%

FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Models

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

机构 * School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦女王学院电子工程与计算机科学学院) Department of Knowledge Technologies, Jožef Stefan Institute(Jožef Stefan研究所知识技术系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏