arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-14 至 2025-11-14 共收录 60 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 5 篇

2501.18638 2025-11-14 cs.CR cs.AI cs.CL 62%

Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation

Daniel Schwartz, Dmitriy Bespalov, Zhe Wang, Ninad Kulkarni, Yanjun Qi

机构 * Amazon Bedrock Science(亚马逊Bedrock科学) Drexel University(德雷塞尔大学) University of Virginia(弗吉尼亚大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 14 pages, 5 figures; published in EMNLP 2025 ; Code at: https://github.com/dsbuddy/GAP-LLM-Safety

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态评测 12 篇

2410.17337 2025-11-14 cs.CL cs.AI cs.IR 84%

Captions Speak Louder than Images: Generalizing Foundation Models for E-commerce from High-quality Multimodal Instruction Data

Xinyi Ling, Hanwen Du, Bo Peng, Zhihui Zhu, Xia Ning

机构 * Department of Computer Science and Engineering, The Ohio State University(计算机科学与工程系,俄亥俄州立大学) Translational Data Analytics Institute, The Ohio State University(转化数据分析研究所,俄亥俄州立大学) Department of Biomedical Informatics, The Ohio State University(生物医学信息学系,俄亥俄州立大学)

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CL、cs.AI

Comments IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10075 2025-11-14 cs.CL 83%

Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts

Xanh Ho, Yun-Ang Wu, Sunisth Kumar, Florian Boudin, Atsuhiro Takasu, Akiko Aizawa

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21572 2025-11-14 cs.CL 83%

Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling

Shengwu. Xiong, Tianyu. Zou, Cong. Wang, Xuelong Li

机构 * Interdisciplinary Artificial Intelligence Research Institute, Wuhan College(交叉学科人工智能研究 institute,武汉学院) School of Computer and Artificial Intelligence, Wuhan University of Technology(计算机与人工智能学院,武汉理工大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Sanya Science and Education Innovation Park, Wuhan University of Technology(三亚科学教育创新园,武汉理工大学) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) School of Mathematics and Statistics, Northwestern Polytechnical University(数学与统计学院,西北工业大学) Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究所(TeleAI))

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CL

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07250 2025-11-14 cs.CV cs.AI 81%

MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs

Tianhao Peng, Haochen Wang, Yuanxing Zhang, Zekun Wang, Zili Wang, Gavin Chang, Jian Yang, Shihao Li, Yanghai Wang, Xintao Wang, Houyi Li, Wei Ji, Pengfei Wan, Steven Huang, Zhaoxiang Zhang, Jiaheng Liu

机构 * Nanjing University(南京大学) CASIA(中国科学院自动化研究所) Kuaishou Technology(快手科技) M-A-P

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref The Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15721 2025-11-14 cs.CL cs.AI 81%

EcomMMMU: Strategic Utilization of Visuals for Robust Multimodal E-commerce Models

Xinyi Ling, Hanwen Du, Zhihui Zhu, Xia Ning

机构 * Department of Computer Science and Engineering, The Ohio State University(俄亥俄州立大学计算机科学与工程系) Translational Data Analytics Institute, The Ohio State University(俄亥俄州立大学转化数据分析研究所) Department of Biomedical Informatics, The Ohio State University(俄亥俄州立大学生物医学信息学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments ICJNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09581 2025-11-14 eess.IV q-bio.QM 78%

Clinically-aligned Multi-modal Chest X-ray Classification

Phillip Sloan, Edwin Simpson, Majid Mirmehdi

专题命中 多模态评测 :multi-modal(title);multimodal(abstract)

Comments 9 Pages, 2 Figures, 3 Tables & 2 Supplementary Tables in Appendix. Accepted to ML4H 2025 (Proceedings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16711 2025-11-14 cs.RO cs.CV cs.LG 74%

Depth Matters: Multimodal RGB-D Perception for Robust Autonomous Agents

Mihaela-Larisa Clement, Mónika Farsang, Felix Resch, Mihai-Teodor Stanusoiu, Radu Grosu

机构 * CPS, Technische Universität Wien (TU Wien)(CPS,维也纳技术大学)

专题命中 多模态评测 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26125 2025-11-14 cs.CV cs.AI 73%

WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios

Runsheng Xu, Hubert Lin, Wonseok Jeon, Hao Feng, Yuliang Zou, Liting Sun, John Gorman, Ekaterina Tolstaya, Sarah Tang, Brandyn White, Ben Sapp, Mingxing Tan, Jyh-Jing Hwang, Dragomir Anguelov

机构 * Waymo LLC(Waymo公司)

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10027 2025-11-14 cs.CL cs.AI 62%

LLMCARE: early detection of cognitive impairment via transformer models enhanced by LLM-generated synthetic data

Ali Zolnour, Hossein Azadmaleki, Yasaman Haghbin, Fatemeh Taherinezhad, Mohamad Javad Momeni Nezhad, Sina Rashidi, Masoud Khani, AmirSajjad Taleban, Samin Mahdizadeh Sani, Maryam Dadkhah, James M. Noble, Suzanne Bakken, Yadollah Yaghoobzadeh, Abdol-Hossein Vahabie, Masoud Rouhizadeh, Maryam Zolnoori

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08215 2025-11-14 cs.CV cs.LG 57%

Evaluating Gemini LLM in Food Image-Based Recipe and Nutrition Description with EfficientNet-B4 Visual Backbone

Rizal Khoirul Anam

机构 * Department of Computer Science and Technology(计算机科学与技术系) Nanjing University of Information Science and Technology(南京信息工程大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12791 2025-11-14 cs.CV eess.IV 57%

Mitigating Perception Bias: A Training-Free Approach to Enhance LMM for Image Quality Assessment

Baoliang Chen, Siyi Pan, Dongxu Wu, Liang Xie, Xiangjie Sui, Lingyu Zhu, Hanwei Zhu

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10418 2025-11-14 cs.DB 50%

CityVerse: A Unified Data Platform for Multi-Task Urban Computing with Large Language Models

Yaqiao Zhu, Hongkai Wen, Mark Birkin, Man Luo

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态Agent 3 篇

2511.10017 2025-11-14 cs.CV 83%

AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models

Xinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu, Zhen Li, Na Zhao

机构 * University of Science and Technology of China(科学技术大学) Singapore University of Technology and Design(新加坡科技设计大学) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09804 2025-11-14 cs.AI 79%

SlideBot: A Multi-Agent Framework for Generating Informative, Reliable, Multi-Modal Presentations

Eric Xie, Danielle Waterfield, Michael Kennedy, Aidong Zhang

专题命中 多模态Agent :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments 32 pages, 14 figures, accepted into EAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09727 2025-11-14 cs.RO cs.AI cs.LG 57%

Baby Sophia: A Developmental Approach to Self-Exploration through Self-Touch and Hand Regard

Stelios Zarifis, Ioannis Chalkiadakis, Artemis Chardouveli, Vasiliki Moutzouri, Aggelos Sotirchos, Katerina Papadimitriou, Panagiotis Filntisis, Niki Efthymiou, Petros Maragos, Katerina Pastra

机构 * Robotics Institute, Athena Research Center(机器人研究所,阿提卡研究中心) HERON – Hellenic Robotics Center of Excellence(HERON – 希腊机器人 excellence 中心) Institute for Language and Speech Processing, Athena Research Center(语言和语音处理研究所,阿提卡研究中心)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 5 pages, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态训练与对齐 8 篇

2511.10035 2025-11-14 cs.CV 79%

DGFusion: Dual-guided Fusion for Robust Multi-Modal 3D Object Detection

Feiyang Jia, Caiyan Jia, Ailin Liu, Shaoqing Xu, Qiming Xia, Lin Liu, Lei Yang, Yan Gong, Ziying Song

机构 * School of Computer Science and Technology, Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, Beijing Jiaotong University(计算机科学与技术学院、交通数据挖掘与具身智能北京市重点实验室、北京交通大学) State Key Laboratory of Internet of Things for Smart City and Department of Electrome chanical Engineering, University of Macau(智能城市物联网国家重点实验室、澳门大学机电工程系) Fujian Key Laboratory of Sensing and Computing for Smart Cities, Xiamen University(智能城市感知与计算福建省重点实验室、厦门大学) School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院、南洋理工大学) State Key Laboratory of Robotics and System, Harbin Institute of Technology(机器人系统国家重点实验室、哈尔滨工业大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12174 2025-11-14 cs.CV cs.RO 79%

UniGS: Unified Geometry-Aware Gaussian Splatting for Multimodal Rendering

Yusen Xie, Zhenmin Huang, Jianhao Jiao, Dimitrios Kanoulas, Jun Ma

机构 * HKUST (GZ)(香港科技大学(广州)) HKUST(香港科技大学) UCL(伦敦大学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10081 2025-11-14 cs.CV 70%

GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs

Yuxiang Duan, Ao Li, Yingqin Li, Luyu Li, Pengwei Wang

机构 * Shandong University(山东大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10211 2025-11-14 cs.CV 57%

HeatV2X: Scalable Heterogeneous Collaborative Perception via Efficient Alignment and Interaction

Yueran Zhao, Zhang Zhang, Chao Sun, Tianze Wang, Chao Yue, Nuoran Li

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10087 2025-11-14 cs.RO cs.AI cs.LG 57%

Opinion: Towards Unified Expressive Policy Optimization for Robust Robot Learning

Haidong Huang, Haiyue Zhu. Jiayu Song, Xixin Zhao, Yaohua Zhou, Jiayi Zhang, Yuze Zhai, Xiaocong Li

机构 * Eastern Institute of Technology(东部技术研究所) University of Nottingham(诺丁汉大学) SIMTech, Agency for Science, Technology and Research (A*STAR)(SIMTech,科技研究局(A*STAR)) Southern University of Science and Technology(南方科技大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025 Workshop on Embodied World Models for Decision Making

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08903 2025-11-14 cs.CV 57%

LLM-Guided Probabilistic Fusion for Label-Efficient Document Layout Analysis

Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

机构 * Department of Computer Science, Iowa State University(计算机科学系,爱荷华州立大学) Department of Civil, Construction & Environmental Engineering, Iowa State University(土木、建设与环境工程系,爱荷华州立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09725 2025-11-14 physics.plasm-ph cs.LG physics.data-an 50%

The Data Fusion Labeler (dFL): Challenges and Solutions to Data Harmonization, Labeling, and Provenance in Fusion Energy

Craig Michoski, Matthew Waller, Brian Sammuli, Zeyu Li, Tapan Ganatma Nakkina, Raffi Nazikian, Sterling Smith, David Orozco, Dongyang Kuang, Martin Foltin, Erik Olofsson, Mike Fredrickson, Jerry Louis-Jeune, David R. Hatch, Todd A. Oliver, Mitchell Clark, Steph-Yves Louis

机构 * University of Texas, Austin, TX, USA(德克萨斯大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02909 2025-11-14 cs.LG 50%

Fine-grained Token Allocation Via Operation Pruning for Efficient MLLMs

Aoming Liu, Reuben Tan, Boqing Gong, Bryan A. Plummer

机构 * Boston University(波士顿大学) Microsoft Research(微软研究院)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 其他多模态 6 篇

2511.10648 2025-11-14 cs.CV 57%

Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling

Jiahao Wang, Weiye Xu, Aijun Yang, Wengang Zhou, Lewei Lu, Houqiang Li, Xiaohua Wang, Jinguo Zhu

机构 * Xi’an Jiaotong University(西安交通大学) University of Science and Technology of China(中国科学技术大学) SenseTime Research(商汤科技研究院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025 (The Thirty-Ninth Annual Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09944 2025-11-14 cs.CV 57%

TSPE-GS: Probabilistic Depth Extraction for Semi-Transparent Surface Reconstruction via 3D Gaussian Splatting

Zhiyuan Xu, Nan Min, Yuhang Guo, Tong Wei

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments AAAI26 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05162 2025-11-14 cs.CY cs.AI 57%

Artificial-Intelligence Grading Assistance for Handwritten Components of a Calculus Exam

Gerd Kortemeyer, Alexander Caspar, Daria Horica

机构 * Michigan State University(密歇根州立大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10448 2025-11-14 cs.RO 50%

Improving dependability in robotized bolting operations

Lorenzo Pagliara, Violeta Redondo, Enrico Ferrentino, Manuel Ferre, Pasquale Chiacchio

机构 * Department of Information Engineering, Electrical Engineering and Applied Mathematics (DIEM), University of Salerno(信息工程、电气工程与应用数学系,萨勒诺大学) Universidad Politécnica de Madrid, Centre for Automation and Robotics (CAR) UPM-CSIC(马德里理工大学,自动化与机器人中心)

专题命中 其他多模态 :multimodal(abstract)

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10185 2025-11-14 cond-mat.mtrl-sci 50%

Dual-Mode Luminescent Thermometry in LiYO2:Nd3+,Yb3+ Enabled by Structural Phase Transition and Phonon-Assisted Energy Transfer

M. Tahir Abbas, M. Szymczak, D. Szymanski, M. Drozd, G. Chen, L. Marciniak

专题命中 其他多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09655 2025-11-14 astro-ph.IM astro-ph.HE cs.LG 50%

Analysis of the TAIGA-HiSCORE Data Using the Latent Space of Autoencoders

Yu. Yu. Dubenskaya, S. P. Polyakov, A. P. Kryukov, A. P. Demichev, E. O. Gres, E. B. Postnikov, A. Yu. Razumov, P. A. Volchugov, D. P. Zhurov

机构 * Lomonosov Moscow State University, Skobeltsyn Institute of Nuclear Physics(罗蒙诺索夫莫斯科国立大学,斯科贝利辛核物理研究所) Institute for Informatics and Automation Problems of the National Academy of Science of the Republic of Armenia(亚美尼亚国家科学院信息与自动化问题研究所) Research Institute of Applied Physics(应用物理研究 institutes)

专题命中 其他多模态 :multimodal(abstract)

Comments 16 pages, 7 figures, Proceedings of The 9th International Conference on Deep Learning in Computational Physics, July 2-4, 2025, Moscow, Russia

详情

展开后加载摘要…

URL PDF HTML 收藏