arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-24 至 2025-10-24 共收录 45 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 7 篇

2510.20291 2025-10-24 cs.CV cs.AI 82%

A Parameter-Efficient Mixture-of-Experts Framework for Cross-Modal Geo-Localization

LinFeng Li, Jian Zhao, Zepeng Yang, Yuhang Song, Bojun Lin, Tianle Zhang, Yuchen Yuan, Chi Zhang, Xuelong Li

机构 * The Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) East China Normal University(东华师范大学) National Tsing Hua University(国立清华大学)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Journal ref IROS 2025 Robosense Cross-Modal Drone Navigation Challenge first place

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15235 2025-10-24 cs.CV cs.CL 62%

ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding

Jialiang Kang, Han Shu, Wenshuo Li, Yingjie Zhai, Xinghao Chen

机构 * Peking University(北京大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22651 2025-10-24 cs.CV cs.CL cs.LG 62%

Sherlock: Self-Correcting Reasoning in Vision-Language Models

Yi Ding, Ruqi Zhang

机构 * Department of Computer Science, Purdue University, USA(计算机科学系,普渡大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Published at NeurIPS 2025, 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11261 2025-10-24 cs.AI cs.CL 62%

Sycophancy in Vision-Language Models: A Systematic Analysis and an Inference-Time Mitigation Framework

Yunpu Zhao, Rui Zhang, Junbin Xiao, Changxin Ke, Ruibo Hou, Yifan Hao, Ling Li

机构 * School of Computer Science and Technology, University of Science and Technology of China(计算机科学与技术学院,中国科学技术大学) State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences(处理器国家重点实验室,中国科学院计算技术研究所) Department of Computer Science, National University of Singapore(计算机科学系,新加坡国立大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences(软件智能研究中心,中国科学院软件研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Journal ref Neurocomputing, Volume 659, 2026, 131217

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20696 2025-10-24 cs.CV 57%

Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward

Jing Bi, Guangyu Sun, Ali Vosoughi, Chen Chen, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) University of Central Florida(中央佛罗里达大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21401 2025-10-24 cs.CV 57%

JaiLIP: Jailbreaking Vision-Language Models via Loss Guided Image Perturbation

Md Jueal Mia, M. Hadi Amini

机构 * Knight Foundation School of Computing and Information Sciences (KFSCIS)(骑士基金会计算与信息科学学院) Florida International University(佛罗里达国际大学) Sustainability, Optimization, and Learning for InterDependent networks laboratory (solid lab)(可持续性、优化与互依赖网络学习实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20278 2025-10-24 cs.LG 50%

KCM: KAN-Based Collaboration Models Enhance Pretrained Large Models

Guangyu Dai, Siliang Tang, Yueting Zhuang

机构 * Zhejiang University(浙江大学)

专题命中 图文多模态 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 4 篇

2510.20223 2025-10-24 cs.CR cs.MM 83%

Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations

Divyanshu Kumar, Shreyas Jena, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23155 2025-10-24 cs.CV 83%

PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling

Xiao Yu, Yan Fang, Xiaojie Jin, Yao Zhao, Yunchao Wei

机构 * Institute of Information Science, Beijing Jiaotong University(信息科学学院,北京交通大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV

Comments This paper is accepted by 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19671 2025-10-24 cs.SD cs.AI eess.AS 62%

Automated evaluation of children's speech fluency for low-resource languages

Bowen Zhang, Nur Afiqah Abdul Latiff, Justin Kan, Rong Tong, Donny Soh, Xiaoxiao Miao, Ian McLoughlin

机构 * College of Computing \& Data Science

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments 5 pages, 2 figures, conference

Journal ref Proc. Interspeech 2025, pp. 1948-1952, 17-21 Aug. 2025, Rotterdam, The Netherlands

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23203 2025-10-24 cs.RO 50%

CE-Nav: Flow-Guided Reinforcement Refinement for Cross-Embodiment Local Navigation

Kai Yang, Tianlin Zhang, Zhengbo Wang, Zedong Chu, Xiaolong Wu, Yang Cai, Mu Xu

机构 * AMAP, Alibaba Group(阿里集团AMAP)

专题命中 音频语音多模态 :multi-modal(abstract)

Comments Project Page: https://ce-nav.github.io/. Code is available at https://github.com/amap-cvlab/CE-Nav

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 1 篇

2510.20699 2025-10-24 q-fin.CP cs.AI 57%

Fusing Narrative Semantics for Financial Volatility Forecasting

Yaxuan Kong, Yoontae Hwang, Marcus Kaiser, Chris Vryonides, Roel Oomen, Stefan Zohren

机构 * University of Oxford(牛津大学) Pusan National University(釜山国立大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments The 6th ACM International Conference on AI in Finance (ICAIF 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2510.20393 2025-10-24 cs.CV cs.MM 81%

Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval

Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim

机构 * Singapore Management University(新加坡管理大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20193 2025-10-24 cs.IR cs.CL cs.CV cs.LG 81%

Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures

Rahul Raja, Arpita Vats

机构 * Carnegie Mellon University(卡内基梅隆大学) Boston University(波士顿大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments In Proceedings of the 2nd ACM Workshop in AI-powered Question and Answering Systems (AIQAM '25), October 27-28, 2025, Dublin, Ireland. ACM, New York, NY, USA, 8 pages. https://doi.org/10.1145/3746274.3760393

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 4 篇

2510.20637 2025-10-24 cs.LG 78%

Large Multimodal Models-Empowered Task-Oriented Autonomous Communications: Design Methodology and Implementation Challenges

Hyun Jong Yang, Hyunsoo Kim, Hyeonho Noh, Seungnyun Kim, Byonghyo Shim

机构 * Seoul National University(首尔国立大学) Hanbat National University(翰baum国立大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20595 2025-10-24 stat.ML cs.LG 78%

Diffusion Autoencoders with Perceivers for Long, Irregular and Multimodal Astronomical Sequences

Yunyi Shen, Alexander Gagliano

机构 * EECS, MIT(麻省理工学院电子工程与计算机科学系) IAIFI, MIT(麻省理工学院天文研究所) CfA, Harvard(哈佛大学天文台)

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00565 2025-10-24 stat.CO cs.LG math.ST stat.TH 78%

Sampling from multi-modal distributions with polynomial query complexity in fixed dimension via reverse diffusion

Adrien Vacher, Omar Chehab, Anna Korba

机构 * CREST, ENSAE Institut Polytechnique de Paris(CREST,ENSAE 巴黎高等理工学院)

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15382 2025-10-24 cs.LG cs.AI cs.RO 57%

Towards Robust Zero-Shot Reinforcement Learning

Kexin Zheng, Lauriane Teyssier, Yinan Zheng, Yu Luo, Xianyuan Zhan

机构 * The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments Neurips 2025, 29 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 11 篇

2501.01243 2025-10-24 cs.CV cs.AI cs.CL 82%

Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants

Lixiong Qin, Shilong Ou, Miaoxuan Zhang, Jiangning Wei, Yuhang Zhang, Xiaoshuai Song, Yuchen Liu, Mei Wang, Weiran Xu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Normal University(北京师范大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 50 pages, 14 figures, 42 tables. NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20381 2025-10-24 cs.CL cs.AI 81%

VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation

Son T. Luu, Trung Vo, Hiep Nguyen, Khanh Quoc Tran, Kiet Van Nguyen, Vu Tran, Ngan Luu-Thuy Nguyen, Le-Minh Nguyen

机构 * Japan Advanced Institute of Science and Technology(日本先进科学研究院) University of Information Technology(信息技术大学) Vietnam National University(越南国家大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments VLSP 2025 MLQA-TSR Share Task

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19892 2025-10-24 cs.CL cs.AI 81%

Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities

Nishant Balepur, Dang Nguyen, Dayeon Ki

机构 * University of Maryland(马里兰大学)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted as a Spotlight paper at the EMNLP 2025 Wordplay Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20632 2025-10-24 cs.AI 79%

Towards Reliable Evaluation of Large Language Models for Multilingual and Multimodal E-Commerce Applications

Shuyi Xie, Ziqin Liew, Hailing Zhang, Haibo Zhang, Ling Hu, Zhiqiang Zhou, Shuman Liu, Anxiang Zeng

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11520 2025-10-24 cs.CV 79%

mmWalk: Towards Multi-modal Multi-view Walking Assistance

Kedi Ying, Ruiping Liu, Chongyan Chen, Mingzhe Tao, Hao Shi, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen

机构 * CV:HCI, KIT(KIT计算机视觉与人机交互中心) Hunan University(湖南大学) ETH Zurich(苏黎世联邦理工学院) University of Texas at Austin(德克萨斯大学奥斯汀分校) Zhejiang University(浙江大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025 Datasets and Benchmarks Track. Data and Code: https://github.com/KediYing/mmWalk

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04462 2025-10-24 cs.CL cs.AI 62%

Benchmarking GPT-5 for biomedical natural language processing

Yu Hou, Zaifu Zhan, Min Zeng, Yifan Wu, Shuang Zhou, Rui Zhang

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20612 2025-10-24 cs.CV cs.CL cs.LG 62%

Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models

Peter Robicheaux, Matvei Popov, Anish Madan, Isaac Robinson, Joseph Nelson, Deva Ramanan, Neehar Peri

机构 * Roboflow Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments The first two authors contributed equally. This work has been accepted to the Neural Information Processing Systems (NeurIPS) 2025 Datasets & Benchmark Track. Project Page: https://rf100-vl.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16793 2025-10-24 cs.CV 57%

REOBench: Benchmarking Robustness of Earth Observation Foundation Models

Xiang Li, Yong Tao, Siyuan Zhang, Siwei Liu, Zhitong Xiong, Chunbo Luo, Lu Liu, Mykola Pechenizkiy, Xiao Xiang Zhu, Tianjin Huang

机构 * University of Bristol, UK(英国布里斯托大学) University of Exeter, UK(英国埃克塞特大学) South China Normal University, China(华南师范大学) The University of Aberdeen, UK(英国阿伯丁大学) Technical University of Munich, Germany(慕尼黑技术大学) Eindhoven University of Technology, NL(埃因霍温理工大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted to NeruIPS 2025 D&B Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15560 2025-10-24 cs.CV cs.HC eess.IV eess.SP 57%

QUB-PHEO: A Visual-Based Dyadic Multi-View Dataset for Intention Inference in Collaborative Assembly

Samuel Adebayo, Seán McLoone, Joost C. Dessing

机构 * Centre for Intelligent Autonomous Manufacturing Systems, Queen’s University Belfast(智能自主制造系统研究中心,女王大学贝尔法斯特) School of Electronics, Electrical Engineering and Computer Science, Queen’s University Belfast(电子、电气工程与计算机科学学院,女王大学贝尔法斯特) School of Psychology, Queen’s University Belfast(心理学学院,女王大学贝尔法斯特)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Journal ref IEEE Access, Vol. 12, pp. 157050-157066, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06259 2025-10-24 cs.CY cs.LG 50%

Beyond Static Knowledge Messengers: Towards Adaptive, Fair, and Scalable Federated Learning for Medical AI

Jahidul Arafat, Fariha Tasmin, Sanjaya Poudel, Iftekhar Haider

机构 * Department of Computer Science and Software Engineering, Auburn University(计算机科学与软件工程系,阿伯丁大学) Department of Information and Communication Technology, Bangladesh University of Professionals(信息与通信技术系,孟加拉国专业大学) Mymensingh Medical College and Hospital(迈明辛医疗学院和医院)

专题命中 多模态评测 :multi-modal(abstract)

Comments 20 pages, 4 figures, 14 tables. Proposes Adaptive Fair Federated Learning (AFFL) algorithm and MedFedBench benchmark suite for healthcare federated learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12816 2025-10-24 quant-ph 50%

Floquet Analysis of Frequency Collisions

Kentaro Heya, Moein Malekakhlagh, Seth Merkel, Naoki Kanazawa, Emily Pritchett

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态Agent 2 篇

2510.20409 2025-10-24 cs.HC 50%

Designing Intent Communication for Agent-Human Collaboration

Yi Li, Francesco Chiossi, Helena Anna Frijns, Jan Leusmann, Julian Rasch, Robin Welsch, Philipp Wintersberger, Florian Michahelles, Albrecht Schmidt

专题命中 多模态Agent :multi-modal(abstract)

Journal ref 24th International Conference on Mobile and Ubiquitous Multimedia - December 01--04, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏