arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-22 至 2025-10-22 共收录 50 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 3 篇

2505.17098 2025-10-22 cs.CL cs.CV 81%

TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence Configuration

Yanshu Li, Jianjiang Yang, Tian Yun, Pinyuan Feng, Jinfa Huang, Ruixiang Tang

机构 * Brown University(布朗大学) University of Bristol(布里斯托大学) Columbia University(哥伦比亚大学) University of Rochester(罗切斯特大学) Rutgers University(罗格斯大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments EMNLP2025 Main, 28 pages, 11 figures, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17800 2025-10-22 cs.CV cs.CL cs.LG 62%

Glyph: Scaling Context Windows via Visual-Text Compression

Jiale Cheng, Yusen Liu, Xinyu Zhang, Yulin Fei, Wenyi Hong, Ruiliang Lyu, Weihan Wang, Zhe Su, Xiaotao Gu, Xiao Liu, Yushi Bai, Jie Tang, Hongning Wang, Minlie Huang

机构 * The Conversational Artificial Intelligence (CoAI) Group, Tsinghua University(清华大学对话人工智能(CoAI)小组) Zhipu AI(智谱AI) The Knowledge Engineering Group (KEG), Tsinghua University(清华大学知识工程小组(KEG))

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18262 2025-10-22 cs.CV 57%

UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding

Da Zhang, Chenggang Rong, Bingyu Li, Feiyu Wang, Zhiyuan Zhao, Junyu Gao, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI)、中国电信)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments We have released V1, which only reports the test results. Our work is still ongoing, and the next version will be coming soon

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 3 篇

2505.03739 2025-10-22 cs.CL cs.AI 84%

VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model

Zuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao, Lijiang Li, Peixian Chen, Mengdan Zhang, Hang Shao, Jian Li, Jinlong Peng, Haoyu Cao, Ke Li, Rongrong Ji, Xing Sun

机构 * Tencent Youtu Lab(腾讯优图实验室) Nanjing University(南京大学) Xiamen University(厦门大学)

专题命中 音频语音多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CL、cs.AI;MLLM(comments)

Comments Training and Inference Codes: https://github.com/VITA-MLLM/VITA-Audio

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16122 2025-10-22 cs.CL 83%

Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection

Ankan Mullick, Saransh Sharma, Abhik Jana, Pawan Goyal

机构 * IIT Kharagpur, India(印度理工学院Kharagpur) Adobe Research, India(Adobe研究) IIT Bhubaneswar, India(印度理工学院Bhubaneswar)

专题命中 音频语音多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

Comments EMNLP 2025 Main Conference Full Paper

Journal ref EMNLP 2025 Main Conference Full Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02236 2025-10-22 cs.CV cs.MM cs.SD eess.AS 82%

3D Audio-Visual Segmentation

Artem Sokolov, Swapnil Bhosale, Xiatian Zhu

机构 * University of Surrey, UK(Surrey大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted at the NeurIPS 2024 Workshop on Audio Imagination; this version updates the project page link

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2510.17305 2025-10-22 cs.CV cs.MM 84%

LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding

ZhaoYang Han, Qihan Lin, Hao Liang, Bowen Chen, Zhou Liu, Wentao Zhang

机构 * Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学)

专题命中 视频多模态 :omni-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments Submitted to ARR Rolling Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18411 2025-10-22 cs.CL cs.LG 79%

DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding

Yue Jiang, Jichu Li, Yang Liu, Dingkang Yang, Feng Zhou, Quyu Kong

机构 * Fudan University(复旦大学) Center for Applied Statistics and School of Statistics, Renmin University of China(应用统计中心和中国人民大学统计学院) Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(北京未来区块链与隐私计算高级创新中心)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted by Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18726 2025-10-22 cs.CV 57%

IF-VidCap: Can Video Caption Models Follow Instructions?

Shihao Li, Yuanxing Zhang, Jiangtao Wu, Zhide Lei, Yiwen He, Runzhe Wen, Chenxi Liao, Chengkang Jiang, An Ping, Shuo Gao, Suhan Wang, Zhaozhou Bian, Zijun Zhou, Jingyi Xie, Jiayi Zhou, Jing Wang, Yifan Yao, Weihao Xie, Yingshui Tan, Yanghai Wang, Qianqian Xie, Zhaoxiang Zhang, Jiaheng Liu

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments https://github.com/NJU-LINK/IF-VidCap

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02175 2025-10-22 cs.RO cs.CV cs.LG 57%

VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching

Siyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu, Tao Huang, Chang Xu

机构 * School of Computer Science, University of Sydney(悉尼大学计算机科学学院) John Hopcropt Center for Computer Science, Shanghai Jiao Tong University(上海交通大学约翰·霍普克罗夫特计算机科学中心)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22205 2025-10-22 cs.RO 50%

From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment

Ke Ye, Jiaming Zhou, Yuanfeng Qiu, Jiayi Liu, Shihui Zhou, Kun-Yu Lin, Junwei Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The University of Hong Kong(香港大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 视频多模态 :multimodal(abstract)

Comments More details and videos can be found at: https://yipko.com/super-mimic

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 5 篇

2510.18703 2025-10-22 cs.CV 85%

Exploring a Unified Vision-Centric Contrastive Alternatives on Multi-Modal Web Documents

Yiqi Lin, Alex Jinpeng Wang, Linjie Li, Zhengyuan Yang, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学展示实验室) Central South University(中南大学) Microsoft(微软公司)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);cross-modal(abstract);image-text(abstract)

Comments Project page: this https://linyq17.github.io/VC2L/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18303 2025-10-22 cs.CV 79%

Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language Models

Lehan Wang, Yi Qin, Honglong Yang, Xiaomeng Li

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04641 2025-10-22 cs.LG math.ST stat.ML stat.TH 78%

A Statistical Theory of Contrastive Pre-training and Multimodal Generative AI

Kazusato Oko, Licong Lin, Yuhang Cai, Song Mei

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17960 2025-10-22 astro-ph.IM astro-ph.CO 75%

AION-1: Omnimodal Foundation Model for Astronomical Sciences

Liam Parker, Francois Lanusse, Jeff Shen, Ollie Liu, Tom Hehir, Leopoldo Sarra, Lucas Meyer, Micah Bowles, Sebastian Wagner-Carena, Helen Qu, Siavash Golkar, Alberto Bietti, Hatim Bourfoune, Nathan Casserau, Pierre Cornette, Keiya Hirashima, Geraud Krawezik, Ruben Ohana, Nicholas Lourie, Michael McCabe, Rudy Morel, Payel Mukhopadhyay, Mariel Pettee, Bruno Regaldo-Saint Blancard, Kyunghyun Cho, Miles Cranmer, Shirley Ho

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract)

Comments Accepted at Neural Information Processing Systems (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04039 2025-10-22 cs.CR cs.AI 57%

BlockScan: Detecting Anomalies in Blockchain Transactions

Jiahao Yu, Xian Wu, Hao Liu, Wenbo Guo, Xinyu Xing

机构 * UC Santa Barbara(加州大学圣芭芭拉分校) Meta AI New York University(纽约大学) sec3 Northwestern University(西北大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 6 篇

2504.18400 2025-10-22 eess.IV cs.AI cs.CV 81%

A Multimodal Deep Learning Approach for White Matter Shape Prediction in Diffusion MRI Tractography

Yui Lo, Yuqian Chen, Dongnan Liu, Leo Zekelman, Jarrett Rushmore, Yogesh Rathi, Nikos Makris, Alexandra J. Golby, Fan Zhang, Weidong Cai, Lauren J. O'Donnell

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Paper accepted to Human Brain Mapping. 25 pages, 3 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17826 2025-10-22 q-bio.BM cs.AI 79%

Speak to a Protein: An Interactive Multimodal Co-Scientist for Protein Analysis

Carles Navarro, Mariona Torrens, Philipp Thölke, Stefan Doerr, Gianni De Fabritiis

机构 * Acellera Labs(Acellera实验室) ICREA, Universitat Pompeu Fabra(ICREA、庞培法布拉大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01480 2025-10-22 cs.CV 70%

Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning

Kaihang Pan, Yang Wu, Wendong Bu, Kai Shen, Juncheng Li, Yingting Wang, Yunfei Li, Siliang Tang, Jun Xiao, Fei Wu, Hang Zhao, Yueting Zhuang

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02048 2025-10-22 eess.IV cs.AI cs.CV 62%

Regression is all you need for medical image translation

Sebastian Rassmann, David Kügler, Christian Ewert, Martin Reuter

机构 * German Center for Neurodegenerative Diseases (DZNE)(德国神经退行性疾病研究中心)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18345 2025-10-22 cs.CV 57%

GPTFace: Generative Pre-training of Facial-Linguistic Transformer by Span Masking and Weakly Correlated Text-image Data

Yudong Li, Hao Li, Xianxu Hou, Linlin Shen

机构 * Shenzhen University(深圳大学)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments This work was initially drafted in November 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17937 2025-10-22 cs.LG cs.AI 57%

UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts

Fu-Yun Wang, Han Zhang, Michael Gharbi, Hongsheng Li, Taesung Park

机构 * Cuhk(香港中文大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 13 篇

2510.18583 2025-10-22 cs.CV cs.LG 85%

CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text Encoder

Yongmin Lee, Hye Won Chung

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00063 2025-10-22 astro-ph.IM cs.AI 83%

AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy

Jinghang Shi, Xiaoyu Tang, Yang Huang, Yuyang Li, Xiao Kong, Yanxia Zhang, Caizhan Yue

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18014 2025-10-22 cs.CV cs.MM 81%

ManzaiSet: A Multimodal Dataset of Viewer Responses to Japanese Manzai Comedy

Kazuki Kawamura, Kengo Nakai, Jun Rekimoto

机构 * Sony CSL Kyoto(索尼 CSL 京都) The University of Tokyo(东京大学) Yoshimoto Kogyo Holdings Co., Ltd.(吉村工业株式会社)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments ICCV 2025 Workshop on Affective & Behavior Analysis in-the-Wild (ABAW), Honolulu, HI, USA (Oct 19, 2025, HST). 11 pages, 5 figures

Journal ref ICCV 2025 Workshops (ICCVW) / CVF Open Access

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02537 2025-10-22 cs.CV cs.AI 81%

VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning

Hao Yan, Xingchen Liu, Hao Wang, Zhenbiao Cao, Handong Zheng, Liang Yin, Xinxing Su, Zihao Chen, Jihao Wu, Minghui Liao, Chao Weng, Wei Chen, Yuliang Liu, Xiang Bai

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18483 2025-10-22 cs.AI 79%

StarBench: A Turn-Based RPG Benchmark for Agentic Multimodal Decision-Making and Information Seeking

Haoran Zhang, Chenhao Zhu, Sicong Guo, Hanzhe Guo, Haiming Li, Donglin Yu

机构 * University of Michigan(密歇根大学) Stanford University(斯坦福大学) University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18377 2025-10-22 cs.CV 79%

Cross-Modal Scene Semantic Alignment for Image Complexity Assessment

Yuqing Luo, Yixiao Li, Jiang Liu, Jun Fu, Hadi Amirpour, Guanghui Yue, Baoquan Zhao, Padraig Corcoran, Hantao Liu, Wei Zhou

机构 * School of Computer Science Cardiff University(卡迪夫大学计算机科学学院) School of Mathematical Sciences Beihang University(北航数学科学学院) Department of Information Technology University of Klagenfurt(克雷格弗特大学信息技术系) School of Biomedical Engineering Shenzhen University(深圳大学生物医学工程学院) School of Artificial Intelligence Sun Yat-sen University(中山大学人工智能学院)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments 14 pages,2 figures, British Machine Vision Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09507 2025-10-22 cs.AI 79%

An Automated Multi-modal Evaluation Framework for Mobile Intelligent Assistants Based on Large Language Models and Multi-Agent Collaboration

Meiping Wang, Jian Zhong, Rongduo Han, Liming Kang, Zhengkun Shi, Xiao Liang, Xing Lin, Nan Gao, Haining Zhang

机构 * College of Software, Nankai University(南开大学软件学院) vivo AI Lab(vivo人工智能实验室)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18336 2025-10-22 eess.SP 78%

MCANet: A Coherent Multimodal Collaborative Attention Network for Advanced Modulation Recognition in Adverse Noisy Environments

Wangye Jiang, Haoming Yang, Xinyu Lu, Mingyuan Wang, Huimei Sun, Jingya Zhang

专题命中 多模态评测 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏