arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-15 至 2025-10-15 共收录 50 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 3 篇

2510.11852 2025-10-15 cs.LG 80%

Evaluating Open-Source Vision-Language Models for Multimodal Sarcasm Detection

Saroj Basnet, Shafkat Farabi, Tharindu Ranasinghe, Diptesh Kanoji, Marcos Zampieri

机构 * George Mason University(乔治·马歇尔大学) Lancaster University(兰卡斯特大学) University of Surrey(萨里大学)

专题命中 图文多模态 :multimodal(title,abstract)

Comments Accepted to ICDMW 2025 Workshop on Multimodal AI (MMAI). Full workshop info: https://icdmw25mmai.github.io/

Journal ref Proc. IEEE International Conference on Data Mining Workshops (ICDMW 2025), Workshop on Multimodal AI (MMAI 2025), Los Angeles, USA, December 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07214 2025-10-15 cs.CV cs.AI cs.CL 67%

Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions

Akash Ghosh, Arkadeep Acharya, Sriparna Saha, Vinija Jain, Aman Chadha

机构 * Department of Computer Science and Engineering, IIT Patna, India(印度帕纳大学计算机科学与工程系) Stanford University(斯坦福大学) Amazon AI(亚马逊人工智能) Indian Institute of Technology Patna, India(印度帕纳大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments One of the first survey on Visual Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00378 2025-10-15 cs.AI cs.CV 62%

CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding

Shixin Yi, Lin Shang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments The paper is not yet mature and needs further improvement

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 6 篇

2510.11760 2025-10-15 cs.SD cs.AI cs.CV cs.MM 85%

Audio-Guided Visual Perception for Audio-Visual Navigation

Yi Wang, Yinfeng Yu, Fuchun Sun, Liejun Wang, Wendong Zheng

机构 * School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing(丝绸之路多语种认知计算国际联合实验室) Tsinghua University(清华大学) Tianjin University of Technology(天津工业大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Main paper (6 pages). Accepted for publication by International Conference on Virtual Reality and Visualization 2025 (ICVRV 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12133 2025-10-15 cs.CL cs.AI 81%

SafeMT: Multi-turn Safety for Multimodal Language Models

Han Zhu, Juntao Dai, Jiaming Ji, Haoran Li, Chengkun Cai, Pengcheng Wen, Chi-Min Chan, Boyuan Chen, Yaodong Yang, Sirui Han, Yike Guo

机构 * Hong Kong University of Science and Technology(香港理工大学) Peking University(北京大学) University of Edinburgh(爱丁堡大学)

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06689 2025-10-15 cs.SD eess.AS 79%

A Fast and Lightweight Model for Causal Audio-Visual Speech Separation

Wendi Sang, Kai Li, Runxuan Yang, Jianqiang Huang, Xiaolin Hu

机构 * School of Computer Technology and Application(计算机技术与应用学院) Intelligent Computing and Application Laboratory of Qinghai Province(青海省智能计算与应用实验室) Department of Computer Science and Technology(计算机科学与技术系) Institute for AI(人工智能研究院) Tsinghua Laboratory of Brain and Intelligence (THBI)(清华大学脑与智能实验室) IDG/McGovern Institute for Brain Research(IDG/麦克戈维脑研究学院) Tsinghua University(清华大学) Chinese Institute for Brain Research (CIBR)(中国脑科学研究院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted by ECAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11738 2025-10-15 cs.SD cs.AI cs.CV cs.MM 75%

SeeingSounds: Learning Audio-to-Visual Alignment via Text

Simone Carnemolla, Matteo Pennisi, Chiara Russo, Simone Palazzo, Daniela Giordano, Concetto Spampinato

机构 * University of Catania(卡塔尼亚大学)

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.AI、cs.MM

Comments accepted to ACM Multimedia Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11732 2025-10-15 cs.SD cs.AI eess.AS 62%

Serial-Parallel Dual-Path Architecture for Speaking Style Recognition

Guojian Li, Qijie Shao, Zhixian Zhao, Shuiyuan Wang, Zhonghua Fu, Lei Xie

机构 * School of Computer Science, Northwestern Polytechnical University(计算机科学学院,西北工业大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI、eess.AS

Comments Accepted by NCMMSC2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11837 2025-10-15 cs.CR cs.AI 61%

Countermind: A Multi-Layered Security Architecture for Large Language Models

Dominik Schwarz

机构 * Independent Researcher(独立研究者)

专题命中 音频语音多模态 :multimodal(abstract,comments);分类 cs.AI

Comments 33 pages, 3 figures, 6 tables. Keywords: LLM security; defense-in-depth; prompt injection; activation steering; multimodal sandbox; threat modeling

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 7 篇

2509.07447 2025-10-15 cs.CV 79%

In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting

Taiying Peng, Jiacheng Hua, Miao Liu, Feng Lu

机构 * State Key Laboratory of VR Technology and Systems, School of CSE, Beihang University(虚拟现实技术与系统国家重点实验室,北京航空航天大学计算机科学与工程学院) College of AI, Tsinghua University(清华大学人工智能学院)

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12299 2025-10-15 cs.IR 67%

An Empirical Study for Representations of Videos in Video Question Answering via MLLMs

Zhi Li, Yanan Wang, Hao Niu, Julio Vizcarra, Masato Taya

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract)

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16845 2025-10-15 cs.CV cs.AI cs.LG 62%

NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows

Denis Tarasov, Alexander Nikulin, Ilya Zisman, Albina Klepach, Nikita Lyubaykin, Andrei Polubarov, Alexander Derevyagin, Vladislav Kurenkov

机构 * ETH Zürich(苏黎世联邦理工学院) MIPT(莫斯科国立信息安全大学) Skoltech(斯克里普斯基技术大学) Innopolis University(因诺波利斯大学) HSE(俄罗斯高等经济学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments https://github.com/dunnolab/NinA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12483 2025-10-15 cs.RO cs.CV 57%

Fast Visuomotor Policy for Robotic Manipulation

Jingkai Jia, Tong Yang, Xueyao Chen, Chenhuan Liu, Wenqiang Zhang

机构 * Fudan University(复旦大学) MEGVII Technology(MEGVII科技)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12185 2025-10-15 cs.CL cs.SD 57%

Not in Sync: Unveiling Temporal Bias in Audio Chat Models

Jiayu Yao, Shenghua Liu, Yiwei Wang, Rundong Cheng, Lingrui Mei, Baolong Bi, Zhen Xiong, Xueqi Cheng

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of California, Merced(加州大学默塞德分校) Beijing University of Posts and Telecommunications(北京邮电大学) University of Southern California(南加州大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07984 2025-10-15 cs.CV 57%

OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding

Jingli Lin, Chenming Zhu, Runsen Xu, Xiaohan Mao, Xihui Liu, Tai Wang, Jiangmiao Pang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 30 pages, a benchmark designed to evaluate Online Spatio-Temporal understanding from the perspective of an agent actively exploring a scene. Project Page: https://rbler1234.github.io/OSTBench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12434 2025-10-15 cs.CV 57%

VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning

Qi Wang, Yanrui Yu, Ye Yuan, Rui Mao, Tianfei Zhou

机构 * Beijing Institute of Technology(北京理工大学) Shenzhen University(深圳大学)

专题命中 视频多模态 :MLLM(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025. Code: https://github.com/QiWang98/VideoRFT

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 5 篇

2510.12801 2025-10-15 cs.CV cs.IR 79%

DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search

Kartik Narayan, Yang Xu, Tian Cao, Kavya Nerella, Vishal M. Patel, Navid Shiee, Peter Grasch, Chao Jia, Yinfei Yang, Zhe Gan

机构 * Johns Hopkins University(约翰霍普金斯大学) Apple(苹果公司)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15877 2025-10-15 cs.CV cs.CL cs.LG 73%

Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval

Siting Li, Xiang Gao, Simon Shaolei Du

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments NeurIPS 2025; 27 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12323 2025-10-15 cs.AI 70%

RAG-Anything: All-in-One RAG Framework

Zirui Guo, Xubin Ren, Lingrui Xu, Jiahao Zhang, Chao Huang

机构 * The University of Hong Kong(香港大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12474 2025-10-15 cs.CL cs.LG 57%

SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression

Biao Zhang, Lixin Chen, Tong Liu, Bo Zheng

机构 * Taobao & Tmall Group of Alibaba(淘宝与天猫集团)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments Accepted by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12213 2025-10-15 astro-ph.IM 50%

Encapsulating Textual Contents into a MOC data Structure for Advanced Applications

Giuseppe Greco, Thomas Boch, Pierre Fernique, Manon Marchand, Mark Allen, Francois Xavier Pineau, Matthieu Baumann, Marco Molinaro, Roberto De Pietri, Marica Branchesi, Steven Schramm, Gergely Dalya, Elahe Khalouei, Barbara Patricelli, Giulia Stratta

专题命中 跨模态检索 :multimodal(abstract)

Comments Published in Astronomy and Computing; 11 pages, 4 figures

Journal ref Astron. Comput. 54 (2026) 101014

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 10 篇

2510.12254 2025-10-15 cs.LG 82%

FedMMKT:Co-Enhancing a Server Text-to-Image Model and Client Task Models in Multi-Modal Federated Learning

Ningxin He, Yang Liu, Wei Sun, Xiaozhou Ye, Ye Ouyang, Tiegang Gao, Zehui Zhang

机构 * Institute for AI Industry Research, Tsinghua University(人工智能产业研究院,清华大学) School of Software Engineering, Nankai University(软件工程学院,南开大学) Department of Computing, Hong Kong Polytechnic University(计算机学院,香港理工大学) AsiaInfo Technologies(亚信息科技) China-Austria Belt and Road Joint Laboratory on Artificial Intelligence and Advanced Manufacturing, Hangzhou Dianzi University(人工智能与先进制造联合实验室,杭州电子科技大学)

专题命中 多模态生成 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01483 2025-10-15 cs.GR cs.CV 79%

GarmageNet: A Multimodal Generative Framework for Sewing Pattern Design and Generic Garment Modeling

Siran Li, Ruiyang Liu, Chen Liu, Zhendong Wang, Gaofeng He, Yong-Lu Li, Xiaogang Jin, Huamin Wang

机构 * Zhejiang Sci-Tech University Style3D Research Hangzhou China State Key Lab of CAD\&CG, Zhejiang University Shanghai Jiao Tong University Shanghai China State Key Lab of CAD\&CG, Zhejiang University Hangzhou China Style3D Research Shanghai Jiao Tong University

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments 23 pages,20 figures

Journal ref ACM Trans. Graph. 44, 6, Article 216 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12789 2025-10-15 cs.CV cs.AI cs.LG 73%

UniFusion: Vision-Language Model as Unified Encoder in Image Generation

Kevin Li, Manuel Brack, Sudeep Katakol, Hareesh Ravi, Ajinkya Kale

机构 * Adobe Applied Research(Adobe应用研究)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Project page at https://thekevinli.github.io/unifusion/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12325 2025-10-15 cs.IR cs.AI 70%

Causal Inspired Multi Modal Recommendation

Jie Yang, Chenyang Gu, Zixuan Liu

机构 * National University of Singapore Master of Industrial and Systems Engineering(新加坡国立大学工业与系统工程硕士) East China Normal University Master of Library and Information Science(华东师范大学图书馆与信息科学硕士) Shandong Normal University(山东师范大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12000 2025-10-15 cs.SD cs.CL cs.LG 70%

UALM: Unified Audio Language Model for Understanding, Generation and Reasoning

Jinchuan Tian, Sang-gil Lee, Zhifeng Kong, Sreyan Ghosh, Arushi Goel, Chao-Han Huck Yang, Wenliang Dai, Zihan Liu, Hanrong Ye, Shinji Watanabe, Mohammad Shoeybi, Bryan Catanzaro, Rafael Valle, Wei Ping

机构 * CMU(卡内基梅隆大学) NVIDIA(英伟达) UMD(马里兰大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12777 2025-10-15 cs.CV 57%

What If : Understanding Motion Through Sparse Interactions

Stefan Andreas Baumann, Nick Stracke, Timy Phan, Björn Ommer

机构 * Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Project page and code: https://compvis.github.io/flow-poke-transformer

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10742 2025-10-15 cs.AI 57%

The Philosophical Foundations of Growing AI Like A Child

Dezhi Luo, Yijiang Li, Hokin Deng

机构 * University of Michigan(密歇根大学) University of California San Diego(加州大学圣地亚哥分校) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12253 2025-10-15 cs.LG cs.AI 57%

Diffusion Models for Reinforcement Learning: Foundations, Taxonomy, and Development

Changfu Xu, Jianxiong Guo, Yuzhu Liang, Haiyang Huang, Haodong Zou, Xi Zheng, Shui Yu, Xiaowen Chu, Jiannong Cao, Tian Wang

机构 * Jiangxi University of Finance and Economics(江西财经大学) Beijing Normal University(北京师范大学) Anhui University(安徽大学) Macquarie University(麦考瑞大学) University of Technology Sydney(悉尼大学) The University of Science and Technology (Guangzhou)(广州大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12095 2025-10-15 cs.CV 57%

IL3D: A Large-Scale Indoor Layout Dataset for LLM-Driven 3D Scene Generation

Wenxu Zhou, Kaixuan Nie, Hang Du, Dong Yin, Wei Huang, Siqiang Guo, Xiaobo Zhang, Pengbo Hu

机构 * University of Science and Technology of China(中国科学技术大学) Songying Technology(宋英科技)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 9 pages main paper; 15 pages references and appendix

详情

展开后加载摘要…

URL PDF HTML 收藏