arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-27 至 2025-10-27 共收录 56 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 3 篇

2510.20350 2025-10-27 cs.CY cs.AI 57%

What Do AI-Generated Images Want?

Amanda Wasielewski

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10881 2025-10-27 cs.LG 50%

Prior-Guided Diffusion Planning for Offline Reinforcement Learning

Donghyeon Ki, JunHyeok Oh, Seong-Woong Shim, Byung-Jun Lee

机构 * Korea University(韩国大学) Gauss Labs Inc.(Gauss实验室)

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态评测 11 篇

2306.13394 2025-10-27 cs.CV 84%

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, Rongrong Ji, Caifeng Shan, Ran He

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院) Tencent Youtu Lab(腾讯优图实验室) Xiamen University(厦门大学) CASIA(中国科学院自动化研究所)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments NeurIPS DB 2025 Spotlight, Project Page: https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21603 2025-10-27 cs.IR cs.CL 83%

Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep Research

Kuicai Dong, Shurui Huang, Fangda Ye, Wei Han, Zhi Zhang, Dexun Li, Wenjun Li, Qu Yang, Gang Wang, Yichao Wang, Chen Zhang, Yong Liu

机构 * Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21111 2025-10-27 cs.CV 83%

PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments

Weijie Zhou, Xuantang Xiong, Yi Peng, Manli Tao, Chaoyang Zhao, Honghui Dong, Ming Tang, Jinqiao Wang

机构 * Beijing Jiaotong University(北京交通大学) Tencent Robotics X & Futian Laboratory, Shenzhen(腾讯机器人X及福田实验室,深圳) Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心,中国科学院自动化研究所) ObjectEye Inc(ObjectEye公司)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments 39th Conference on Neural Information Processing Systemss (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21438 2025-10-27 cs.RO 82%

PREVENT: Proactive Risk Evaluation and Vigilant Execution of Tasks for Mobile Robotic Chemists using Multi-Modal Behavior Trees

Satheeshkumar Veeramani, Zhengxue Zhou, Francisco Munguia-Galeano, Hatem Fakhruldeen, Thomas Roddelkopf, Mohammed Faeik Ruzaij Al-Okby, Kerstin Thurow, Andrew Ian Cooper

机构 * Department of Chemistry and Material Innovation Factory, University of Liverpool(化学系和材料创新工厂,利物浦大学) Center for Life Science Automation, University of Rostock(生命科学自动化中心,罗斯托克大学)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract)

Comments 25 pages, 8 figures, paper submitted to Robotics and Autonomous Systems Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21679 2025-10-27 cs.AI 79%

A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection

Gaku Morio, Harri Rowlands, Dominik Stammbach, Christopher D. Manning, Peter Henderson

机构 * Hitachi, Ltd.(日立公司) Stanford University(斯坦福大学) Centre for the Acceleration of Social Technology(社会技术加速中心) Princeton University(普林斯顿大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Forthcoming in NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11568 2025-10-27 q-bio.QM cs.AI cs.LG 79%

BioCube: A Multimodal Dataset for Biodiversity Research

Stylianos Stasinos, Martino Mensio, Elena Lazovik, Athanasios Trantas

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments published to BiDS'25, 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21214 2025-10-27 cs.CR 78%

Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses

Xingwei Zhong, Kar Wai Fok, Vrizlynn L. L. Thing

专题命中 多模态评测 :MLLM(title);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20967 2025-10-27 cs.CV cs.AI 62%

3DReasonKnee: Advancing Grounded Reasoning in Medical Vision Language Models

Sraavya Sambara, Sung Eun Kim, Xiaoman Zhang, Luyang Luo, Shreya Johri, Mohammed Baharoon, Du Hyun Ro, Pranav Rajpurkar

机构 * Department of Biomedical Informatics, Harvard Medical School(生物医学信息学系,哈佛医学院) Seoul National University Hospital(首尔国立大学医院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21160 2025-10-27 cs.CV 57%

Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study

Guanlin Wu, Boyan Su, Yang Zhao, Pu Wang, Yichen Lin, Hao Frank Yang

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12448 2025-10-27 cs.CV 57%

SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning

Yang Liu, Ming Ma, Xiaomin Yu, Pengxiang Ding, Han Zhao, Mingyang Sun, Siteng Huang, Donglin Wang

机构 * Westlake University(西湖大学) Zhejiang University(浙江大学) Harbin Institute of Technology(哈尔滨工业大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08646 2025-10-27 cs.AI 57%

P-CAFE: Personalized Cost-Aware Incremental Feature Selection For Electronic Health Records

Naama Kashani, Mira Cohen, Uri Shaham

机构 * Bar-Ilan University(巴伊兰大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态Agent 2 篇

2510.20838 2025-10-27 cs.AI cs.MA 70%

Sketch2BIM: A Multi-Agent Human-AI Collaborative Pipeline to Convert Hand-Drawn Floor Plans to 3D BIM

Abir Khan Ratul, Sanjay Acharjee, Somin Park, Md Nazmus Sakib

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12733 2025-10-27 cs.RO cs.AI cs.LG 57%

HYPE: Hybrid Planning with Ego Proposal-Conditioned Predictions

Hang Yu, Julian Jordan, Julian Schmidt, Silvan Lindner, Alessandro Canevaro, Wilhelm Stork

机构 * Mercedes-Benz AG, Research & Development(梅赛德斯-奔驰集团,研发部) Karlsruhe Institute of Technology, ITIV(卡尔斯鲁厄理工学院,ITIV) University of Tübingen(图宾根大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Accepted to IEEE ITSC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态训练与对齐 9 篇

2506.01890 2025-10-27 cs.LG cs.AI 83%

CogniAlign: Word-Level Multimodal Speech Alignment with Gated Cross-Attention for Alzheimer's Detection

David Ortiz-Perez, Manuel Benavent-Lledo, Javier Rodriguez-Juan, Jose Garcia-Rodriguez, David Tomás

机构 * Department of Computer Science and Technology, University of Alicante, Alicante, Spain(计算机科学与技术系,阿利坎特大学,阿利坎特,西班牙) Valencian Graduate School and Research Network of Artificial Intelligence, Valencia, Spain(瓦伦西亚人工智能研究生学校与研究网络,瓦伦西亚,西班牙) Department of Software and Computing Systems, University of Alicante, Alicante, Spain(软件与计算系统系,阿利坎特大学,阿利坎特,西班牙)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Journal ref Knowledge-Based Systems, Vol. 329, 2025, Article 114264

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11194 2025-10-27 cs.CE 82%

Prot2Text-V2: Protein Function Prediction with Multimodal Contrastive Alignment

Xiao Fei, Michail Chatzianastasis, Sarah Almeida Carneiro, Hadi Abdine, Lawrence P. Petalidis, Michalis Vazirgiannis

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments 24 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11128 2025-10-27 cs.LG cs.CV 79%

Lightweight Facial Landmark Detection in Thermal Images via Multi-Level Cross-Modal Knowledge Transfer

Qiyi Tong, Olivia Nocentini, Marta Lagomarsino, Kuanqi Cai, Marta Lorenzini, Arash Ajoudani

机构 * Human-Robot Interfaces and Interaction Laboratory, Istituto Italiano di Tecnologia, Genoa, Italy(人机交互实验室,意大利技术研究院,热那亚,意大利) Ph.D. Program of National Interest in Robotics and Intelligent Machines (DRIM), Università di Genova(机器人与智能机器国家利益博士项目,热那亚大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21551 2025-10-27 cs.LG 78%

Interpretable Multimodal Zero-Shot ECG Diagnosis via Structured Clinical Knowledge Alignment

Jialu Tang, Hung Manh Pham, Ignace De Lathauwer, Henk S. Schipper, Yuan Lu, Dong Ma, Aaqib Saeed

机构 * Eindhoven University of Technology(埃因霍温理工大学) Singapore Management University(新加坡管理大学) Maxima Medical Center(马克斯医疗中心) Erasmus Medical Center(埃因霍温医学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20976 2025-10-27 cs.LG 78%

L^2M^3OF: A Large Language Multimodal Model for Metal-Organic Frameworks

Jiyu Cui, Fang Wu, Haokai Zhao, Minggao Feng, Xenophon Evangelopoulos, Andrew I. Cooper, Yejin Choi

机构 * Department of Chemistry, University of Liverpool(利兹大学化学系) Leverhulme Research Centre for Functional Materials Design, University of Liverpool(利兹大学功能性材料设计研究所以) Department of Computer Science, University of Stanford(斯坦福大学计算机科学系) School of Computer Science and Engineering, University of New South Wales(新南威尔士大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21501 2025-10-27 cs.CV cs.AI 73%

GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs

Guanghao Zheng, Bowen Shi, Mingxing Xu, Ruoyu Sun, Peisen Zhao, Zhibo Zhang, Wenrui Dai, Junni Zou, Hongkai Xiong, Xiaopeng Zhang, Qi Tian

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei Inc.(华为公司)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23579 2025-10-27 cs.LG 71%

BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model

Adibvafa Fallahpour, Andrew Magnuson, Purav Gupta, Shihao Ma, Jack Naimer, Arnav Shah, Haonan Duan, Omar Ibrahim, Hani Goodarzi, Chris J. Maddison, Bo Wang

专题命中 多模态训练与对齐 :multimodal(title)

Comments 28 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21553 2025-10-27 cs.CL cs.LG 57%

Document Understanding, Measurement, and Manipulation Using Category Theory

Jared Claypoole, Yunye Gong, Noson S. Yanofsky, Ajay Divakaran

机构 * SRI International(SRI国际研究院) Brooklyn College(布鲁克林学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23801 2025-10-27 cs.RO 50%

High-Precision Climbing Robot Localization Using Planar Array UWB/GPS/IMU/Barometer Integration

Shuning Zhang, Zhanchen Zhu, Xiangyu Chen, Yunheng Wang, Xu Jiang, Peibo Duan, Renjing Xu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Department of Data Science and AI, Monash University(数据科学与人工智能系,墨尔本大学) Department of Data Science and Artificial Intelligence, Faculty of Infomation Technology, Monash University(数据科学与人工智能系,信息科技学院,墨尔本大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 其他多模态 2 篇

2510.20877 2025-10-27 cs.LG cs.AI 79%

Multimodal Negative Learning

Baoquan Gong, Xiyuan Gao, Pengfei Zhu, Qinghua Hu, Bing Cao

机构 * School of Artificial Intelligence, Tianjin University(人工智能学院,天津大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Comments Published in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17611 2025-10-27 cs.CV 57%

One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection

Jia Guo, Shuai Lu, Lei Fan, Zelin Li, Donglin Di, Yang Song, Weihang Zhang, Wenbing Zhu, Hong Yan, Fang Chen, Huiqi Li, Hongen Liao

机构 * Tsinghua University(清华大学) Beijing Institute of Technology(北京理工大学) Shanghai Jiao Tong University(上海交通大学) City University of Hong Kong(香港城市大学) University of New South Wales(新南威尔士大学) DZ Matrix(DZ矩阵) Fudan University(复旦大学) Rongcheer Co., Ltd.(荣彻科技有限公司)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

Comments Extended version of CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏