arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-05 至 2025-11-05 共收录 54 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 2 篇

2505.21375 2025-11-05 cs.CV 83%

GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution

Fengxiang Wang, Mingshuo Chen, Yueying Li, Di Wang, Haotian Wang, Zonghao Guo, Zefan Wang, Boqi Shan, Long Lan, Yulin Wang, Hongzhen Wang, Wenjing Yang, Bo Du, Jing Zhang

机构 * College of Computer Science and Technology, National University of Defense Technology, China(中国国防科技大学计算机科学与技术学院) Beijing University of Posts and Telecommunications, China(北京邮电大学) School of Computer Science, Wuhan University, China(武汉大学计算机学院) Zhongguancun Academy, China(中关村学院) Tsinghua University, China(清华大学) Beihang University, China(北航大学)

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

Comments NeurlPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25067 2025-11-05 cs.CV 57%

DRIP: Dynamic patch Reduction via Interpretable Pooling

Yusen Peng, Sachin Kumar

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Need more refinement

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 2 篇

2509.19999 2025-11-05 cs.MM cs.CV cs.SD 81%

MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization

Jianxuan Yang, Xiaoran Yang, Lipan Zhang, Xinyue Guo, Zhao Wang, Gongping Huang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08039 2025-11-05 cs.SD cs.CL cs.MM eess.AS 67%

Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning

Shu Wu, Chenxing Li, Wenfu Wang, Hao Zhang, Hualei Wang, Meng Yu, Dong Yu

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.MM、eess.AS

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2507.14809 2025-11-05 cs.CV cs.MM cs.RO 81%

Light Future: Multimodal Action Frame Prediction via InstructPix2Pix

Zesen Zhong, Duomin Zhang, Yijia Li

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen(数据科学学院,香港中文大学(深圳))

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 9 pages including appendix, 4 tables, 8 figures, to be submitted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02182 2025-11-05 cs.CV 79%

Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models

Jinhwan Seo, Yoonki Cho, Junhyug Noh, Sung-eui Yoon

机构 * KAIST(韩国科学技术院) Ewha Womans University(成均馆大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 1st place winner of Grounded Videoqa track at the ICCV2025 Perception Test

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01405 2025-11-05 eess.SP cs.ET 78%

MM-2FSK: Multimodal Frequency Shift Keying for Ultra-Efficient and Robust High-Resolution MIMO Radar Imaging

Vanessa Wirth, Johanna Bräunig, Martin Vossiek, Tim Weyrich, Marc Stamminger

专题命中 视频多模态 :multimodal(title,abstract)

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17664 2025-11-05 cs.CV cs.RO 57%

Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras

Lingdong Kong, Dongyue Lu, Ao Liang, Rong Li, Yuhao Dong, Tianshuai Hu, Lai Xing Ng, Wei Tsang Ooi, Benoit R. Cottereau

机构 * NUS(新加坡国立大学) HKUST(GZ)(香港科技大学(广州)) NTU(南洋理工大学) HKUST(香港科技大学) I 2 R, A*STAR(I2R, A*STAR) IPAL, CNRS IRL 2955, Singapore(IPAL, CNRS IRL 2955, 新加坡) CerCo, CNRS UMR 5549, Université Toulouse III(CerCo, CNRS UMR 5549, 法国图卢兹第三大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025 Spotlight; 43 pages, 17 figures, 16 tables; Project Page at https://talk2event.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02484 2025-11-05 eess.SY cs.SY eess.SP 50%

Using ensemble learning with hybrid graph neural networks and transformers to predict traffic in cities

Ismail Zrigui, Samira Khoulji, Mohamed Larbi Kerkeb

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 4 篇

2511.02358 2025-11-05 cs.CL cs.AI cs.IR cs.LG cs.MM 85%

Let Multimodal Embedders Learn When to Augment Query via Adaptive Query Augmentation

Wongyu Kim, Hochang Lee, Sanghak Lee, Yoonsung Kim, Jaehyun Park

机构 * NC AI(NC人工智能)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted to MMGenSR Workshop (CIKM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02371 2025-11-05 cs.LG 82%

LUMA-RAG: Lifelong Multimodal Agents with Provably Stable Streaming Alignment

Rohan Wandre, Yash Gajewar, Namrata Patel, Vivek Dhalkari

机构 * Dept. of Computer Engineering(计算机工程系) SIES Graduate School of Technology(SIES技术研究生学院) Bharatiya Vidya Bhavan's Sardar Patel Institute of Technology(巴哈里亚·维达·巴万学院萨达尔·帕特尔技术学院)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01892 2025-11-05 cs.LG cs.CL 79%

Retrieval-Augmented Multimodal Depression Detection

Ruibo Hou, Shiyu Teng, Jiaqing Liu, Shurong Chai, Yinhao Li, Lanfen Lin, Yen-Wei Chen

机构 * College of Information Science and Engineering, Ritsumeikan University(信息科学与工程学院,立命馆大学) College of Computer Science and Technology, Zhejiang University(计算机科学与技术学院,浙江大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Accepted in IEEE EMBC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02770 2025-11-05 cs.CL cs.IR 57%

Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval

Hung-Ting Chen, Xiang Liu, Shauli Ravfogel, Eunsol Choi

机构 * Department of Computer Science, New York University(纽约大学计算机科学系) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 8 篇

2510.16888 2025-11-05 cs.CV 83%

Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback

Zongjian Li, Zheyuan Liu, Qihui Zhang, Bin Lin, Feize Wu, Shenghai Yuan, Zhiyuan Yan, Yang Ye, Wangbo Yu, Yuwei Niu, Shaodong Wang, Xinhua Cheng, Li Yuan

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12973 2025-11-05 eess.IV cs.CV cs.LG q-bio.QM 83%

Cross-modal Diffusion Modelling for Super-resolved Spatial Transcriptomics

Xiaofei Wang, Xingxu Huang, Stephen J. Price, Chao Li

机构 * Department of Clinical Neurosciences, University of Cambridge, UK(剑桥大学临床神经科学系) Department of Applied Mathematics and Theoretical Physics, University of Cambridge, UK(剑桥大学应用数学与理论物理系) School of Science and Engineering, University of Dundee, UK(邓迪大学科学与工程学院) School of Medicine, University of Dundee, UK(邓迪大学医学院) Zhejiang Lab, China(浙江实验室)

专题命中 多模态生成 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02046 2025-11-05 cs.CV cs.AI 76%

Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis

Soham Joshi, Shwet Kamal Mishra, Viswanath Gopalakrishnan

机构 * International Institute of Information Technology Bangalore(国际信息科技学院班加罗尔)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.AI

Comments First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02468 2025-11-05 cs.HC cs.CV 70%

HAGI++: Head-Assisted Gaze Imputation and Generation

Chuhan Jiao, Zhiming Hu, Andreas Bulling

机构 * University of Stuttgart(斯图加特大学) The Hong Kong University of Science(香港科学大学)

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Extended version of our UIST'25 paper "HAGI: Head-Assisted Gaze Imputation for Mobile Eye Trackers"

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02423 2025-11-05 eess.SP 67%

LLM4PG: Adapting Large Language Model for Pathloss Map Generation via Synesthesia of Machines

Mingran Sun, Lu Bai, Xiang Cheng, Jianjun Wu

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02091 2025-11-05 cs.LG cs.AI 57%

Natural Building Blocks for Structured World Models: Theory, Evidence, and Scaling

Lancelot Da Costa, Sanjeev Namjoshi, Mohammed Abbas Ansari, Bernhard Schölkopf

机构 * VERSES AI Research Lab(VERSES AI研究实验室) University of Tübingen(图宾根大学) ELLIS Institute, Tübingen(图宾根ELLIS研究所) MPI for Intelligent Systems, Tübingen(图宾根智能系统研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 13 pages, 3 figures, under review for World Modeling Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01894 2025-11-05 cs.GR cs.AI cs.LG 57%

LGCC: Enhancing Flow Matching Based Text-Guided Image Editing with Local Gaussian Coupling and Context Consistency

Fangbing Liu, Pengfei Duan, Wen Li, Yi He

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02530 2025-11-05 cs.AR 50%

Implementation and Evaluation of Stable Diffusion on a General-Purpose CGLA Accelerator

Takuto Ando, Yu Eto, Yasuhiko Nakashima

专题命中 多模态生成 :multi-modal(abstract)

Comments This paper is accepted at 2025 IEEE 18th International Symposium on Embedded Multicore/Many-core Systems-on-Chip (MCSoC)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 16 篇

2511.02234 2025-11-05 cs.MM cs.CL cs.SD 84%

An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM

Jiawei Liu, Enis Berk Çoban, Zarina Schevchenko, Hao Tang, Zhigang Zhu, Michael I Mandel, Johanna Devaney

机构 * The Graduate Center, CUNY(CUNY研究生中心) Brooklyn College, CUNY(CUNY布鲁克林学院) Borough of Manhattan Community College, CUNY(CUNY曼哈顿社区学院) The City College of New York, CUNY(CUNY纽约城市学院)

专题命中 多模态评测 :MLLM(title,abstract);multi-modal(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02778 2025-11-05 cs.CV cs.CL 81%

VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation

Kevin Qinghong Lin, Yuhao Zheng, Hangyu Ran, Dantong Zhu, Dongxing Mao, Linjie Li, Philip Torr, Alex Jinpeng Wang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Project page: https://csu-jpg.github.io/VCode Github: https://github.com/CSU-JPG/VCode

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02495 2025-11-05 cs.CV cs.CL 81%

DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding

Zixuan Liu, Siavash H. Khajavi, Guangkai Jiang

机构 * Department of Computer Science(计算机科学系) Tulane University(Tulane 大学) Department of Industrial Engineering and Management(工业工程与管理系) Aalto University(Aalto 大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Advances in Neural Information Processing Systems 2025 (NeurIPS 2025), Poster, https://neurips.cc/virtual/2025/loc/san-diego/poster/121400

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27195 2025-11-05 cs.CV cs.CL cs.SI 81%

Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions

Caixin Kang, Yifei Huang, Liangyang Ouyang, Mingfang Zhang, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments ICCV2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02794 2025-11-05 cs.AI cs.MA 80%

When One Modality Sabotages the Others: A Diagnostic Lens on Multimodal Reasoning

Chenyu Zhang, Minsol Kim, Shohreh Ghorbani, Jingyao Wu, Rosalind Picard, Patricia Maes, Paul Pu Liang

机构 * Harvard University(哈佛大学) MIT Media Lab(麻省理工学院媒体实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at the Multimodal Algorithmic Reasoning (MAR) Workshop, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15748 2025-11-05 cs.AI 79%

Towards Relaxed Multimodal Inputs for Gait-based Parkinson's Disease Assessment

Minlin Zeng, Zhipeng Zhou, Yang Qiu, Martin J. McKeown, Zhiqi Shen

机构 * College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) Department of Medicine (Neurology), Pacific Parkinsons Research Centre, The University of British Columbia(医学系(神经学),太平洋帕金森研究中心,不列颠哥伦比亚大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22211 2025-11-05 cs.CL 79%

ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding

Kimihiro Hasegawa, Wiradee Imrattanatrai, Zhi-Qi Cheng, Masaki Asada, Susan Holm, Yuran Wang, Ken Fukuda, Teruko Mitamura

机构 * Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学) National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments NAACL2025, Code and Data: https://github.com/kimihiroh/promqa

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20241 2025-11-05 cs.LG cs.AI 79%

DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning

Qi Cao, Ruiyi Wang, Ruiyi Zhang, Sai Ashish Somayajula, Pengtao Xie

机构 * University of California, San Diego(加州大学圣地亚哥分校) Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 28 pages, 10 figures, to appear in NeurIPS 2025 (Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07215 2025-11-05 cs.RO cs.MM 79%

RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation

Feng Yan, Fanfan Liu, Liming Zheng, Yufeng Zhong, Yiyang Huang, Zechao Guan, Chengjian Feng, Lin Ma

机构 * Meituan(美团)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.MM

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision 2025

详情

展开后加载摘要…

URL PDF HTML 收藏