arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46294 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4676 篇

2509.23281 2025-09-30 cs.RO 78%

Preventing Robotic Jailbreaking via Multimodal Domain Adaptation

Francesco Marchiori, Rohan Sinha, Christopher Agia, Alexander Robey, George J. Pappas, Mauro Conti, Marco Pavone

机构 * University of Padova(帕多瓦大学) Stanford University(斯坦福大学) Carnegie Mellon University(卡内基梅隆大学) University of Pennsylvania(宾夕法尼亚大学) Örebro University(奥雷布罗大学) NVIDIA Research(NVIDIA研究)

专题命中 图文多模态 :multimodal(title,abstract)

Comments Project page: https://j-dapt.github.io/. 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19480 2025-09-25 cs.RO cs.LG 78%

OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation

Noriaki Hirose, Catherine Glossop, Dhruv Shah, Sergey Levine

专题命中 图文多模态 :omni-modal(title,abstract)

Comments 9 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18729 2025-09-24 cs.SD 78%

MECap-R1: Emotion-aware Policy with Reinforcement Learning for Multimodal Emotion Captioning

Haoqin Sun, Chenyang Lyu, Xiangyu Kong, Shiwan Zhao, Jiaming Zhou, Hui Wang, Aobo Kong, Jinghua Zhao, Longyue Wang, Weihua Luo, Kaifu Zhang, Yong Qin

机构 * TMCC, College of Computer Science, Nankai University, Tianjin, China(TMCC,计算机科学学院,南开大学,天津,中国) Alibaba International Digital Commerce(阿里巴巴国际数字商务) University of Exeter(埃克塞特大学)

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00449 2025-09-03 cond-mat.mtrl-sci cond-mat.dis-nn cs.LG 78%

Molecular Identification from AFM images using the IUPAC Nomenclature and Attribute Multimodal Recurrent Neural Networks

Jaime Carracedo-Cosme, Carlos Romero-Muñiz, Pablo Pou, Rubén Pérez

专题命中 图文多模态 :multimodal(title,abstract)

Comments 30 pages, 4 figures, 2 tables, includes supplementary information (with additional 21 pages, 9 figures, 1 table)

Journal ref ACS Appl. Mater. Interfaces 15, 22692-22704 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00302 2025-07-22 cs.LG cond-mat.mtrl-sci 78%

Beyond Atomic Geometry Representations in Materials Science: A Human-in-the-Loop Multimodal Framework

Can Polat, Erchin Serpedin, Mustafa Kurban, Hasan Kurban

机构 * Computer Engineering, Texas A\&M University, College Station, TX 77843, USA College of Science Engineering, Hamad Bin Khalifa University, Doha, Qatar Dept. of Electrical \& Computer Engineering, Texas A\&M University at Qatar, Doha, Qatar Orthotics, Ankara University, Ankara, Turkey

专题命中 图文多模态 :multimodal(title,abstract)

Comments Presented at ICML 2025 Workshop on DataWorld

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13624 2025-04-21 eess.SP 78%

PV-VLM: A Multimodal Vision-Language Approach Incorporating Sky Images for Intra-Hour Photovoltaic Power Forecasting

Huapeng Lin, Miao Yu

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03153 2025-04-07 cs.LG 78%

MORAL: A Multimodal Reinforcement Learning Framework for Decision Making in Autonomous Laboratories

Natalie Tirabassi, Sathish A. P. Kumar, Sumit Jha, Arvind Ramanathan

专题命中 图文多模态 :multimodal(title,abstract)

Comments 9 pages, 14 figures and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13164 2025-04-02 cs.LG 78%

VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning

Yongshuo Zong, Ondrej Bohdal, Timothy Hospedales

专题命中 图文多模态 :multimodal(title,abstract)

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08317 2025-03-24 cs.CL cs.AI cs.CV 78%

Mitigating Hallucinations in Multimodal Spatial Relations through Constraint-Aware Prompting

Jiarui Wu, Zhuo Liu, Hangfeng He

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00196 2025-03-05 cs.CV cs.AI cs.CL 78%

DermaSynth: Rich Synthetic Image-Text Pairs Using Open Access Dermatology Datasets

Abdurrahim Yilmaz, Furkan Yuceyalcin, Ece Gokyayla, Donghee Choi, Ozan Erdem, Ali Anil Demircali, Rahmetullah Varol, Ufuk Gorkem Kirabali, Gulsum Gencoglan, Joram M. Posma, Burak Temelkuran

专题命中 图文多模态 :image-text(title);分类 cs.CV、cs.CL、cs.AI

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12662 2025-03-03 cs.CV cs.AI cs.CL 78%

Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models

Shicheng Xu, Liang Pang, Yunchang Zhu, Huawei Shen, Xueqi Cheng

专题命中 图文多模态 :cross-modal(title);分类 cs.CV、cs.CL、cs.AI

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13904 2025-02-14 cs.LG 78%

Privacy-Preserving Personalized Federated Prompt Learning for Multimodal Large Language Models

Linh Tran, Wei Sun, Stacy Patterson, Ana Milanova

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10967 2025-02-13 cs.CV cs.AI cs.CL 78%

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding

Zhanpeng Chen, Mingxiao Li, Ziyang Chen, Nan Du, Xiaolong Li, Yuexian Zou

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07591 2024-12-30 cs.CE 78%

CoinCLIP: A Multimodal Framework for Assessing Viability in Web3 Memecoins

Hou-Wan Long, Hongyang Li, Wei Cai

专题命中 图文多模态 :multimodal(title,abstract)

Comments 4 pages, 1 figure, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15650 2024-12-23 cs.LG 78%

Beyond Human Data: Aligning Multimodal Large Language Models by Iterative Self-Evolution

Wentao Tan, Qiong Cao, Yibing Zhan, Chao Xue, Changxing Ding

专题命中 图文多模态 :multimodal(title,abstract)

Comments AAAI 2025. The code is available at https://github.com/WentaoTan/SENA

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10302 2024-12-16 cs.CV cs.AI cs.CL 78%

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, Zhenda Xie, Yu Wu, Kai Hu, Jiawei Wang, Yaofeng Sun, Yukun Li, Yishi Piao, Kang Guan, Aixin Liu, Xin Xie, Yuxiang You, Kai Dong, Xingkai Yu, Haowei Zhang, Liang Zhao, Yisong Wang, Chong Ruan

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09301 2024-12-11 cond-mat.mtrl-sci physics.comp-ph 78%

Transfer Learning in Materials Informatics: structure-property relationships through minimal but highly informative multimodal input

Dario Massa, Grzegorz Kaszuba, Stefanos Papanikolaou, Piotr Sankowski

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03978 2024-11-07 cs.LG 78%

Customized Multiple Clustering via Multi-Modal Subspace Proxy Learning

Jiawei Yao, Qi Qian, Juhua Hu

专题命中 图文多模态 :multi-modal(title,abstract)

Comments Accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20086 2024-10-01 cs.HC 78%

Optimising EEG decoding with refined sampling and multimodal feature integration

Arash Akbarinia

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15291 2024-09-25 cs.HC cs.CY 78%

Exploring the Feasibility of Multimodal Chatbot AI as Copilot in Pathology Diagnostics: Generalist Model's Pitfall

Mianxin Liu, Jianfeng Wu, Fang Yan, Hongjun Li, Wei Wang, Shaoting Zhang, Zhe Wang

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11629 2024-09-19 cs.IR cs.HC 78%

Designing Interfaces for Multimodal Vector Search Applications

Owen Pendrigh Elliott, Tom Hamer, Jesse Clark

专题命中 图文多模态 :multimodal(title,abstract)

Comments 12 pages, 8 figures, CIKM 2024 MMSR Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21758 2024-08-01 cs.IR 78%

MOSAIC: Multimodal Multistakeholder-aware Visual Art Recommendation

Bereket A. Yilma, Luis A. Leiva

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08173 2024-03-04 cs.LG cs.CR cs.IT math.IT stat.ML 78%

Safeguarding Data in Multimodal AI: A Differentially Private Approach to CLIP Training

Alyssa Huang, Peihan Liu, Ryumei Nakada, Linjun Zhang, Wanrong Zhang

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08531 2023-09-18 cs.CV cs.CL eess.AS eess.IV 78%

Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-training and Multi-modal Tokens

Minsu Kim, Jeongsoo Choi, Soumi Maiti, Jeong Hun Yeo, Shinji Watanabe, Yong Man Ro

专题命中 图文多模态 :multi-modal(title);分类 cs.CV、cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01540 2023-06-05 cs.RO 78%

CLIPGraphs: Multimodal Graph Networks to Infer Object-Room Affinities

Ayush Agrawal, Raghav Arora, Ahana Datta, Snehasis Banerjee, Brojeshwar Bhowmick, Krishna Murthy Jatavallabhula, Mohan Sridharan, Madhava Krishna

专题命中 图文多模态 :multimodal(title,abstract)

Journal ref RO-MAN 2023 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00691 2022-07-05 cs.CY cs.AI cs.CL cs.CV cs.LG 78%

American == White in Multimodal Language-and-Image AI

Robert Wolfe, Aylin Caliskan

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to AI Ethics and Society 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.01871 2019-11-25 cs.HC cs.AI cs.CL cs.CV cs.LG 78%

Help, Anna! Visual Navigation with Natural Multimodal Assistance via Retrospective Curiosity-Encouraging Imitation Learning

Khanh Nguyen, Hal Daumé

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments In EMNLP 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16183 2026-08-25 cs.CV 版本更新 77%

Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation

面向真实世界放射学的专家级视觉-语言基础模型及综合评估

Xiaohong Liu, Guoxing Yang, Yulin Luo, Jiaji Mao, Xiang Zhang, Haibo Wang, Zhiyang He, Ming Gao, Shanghang Zhang, Jun Shen, Guangyu Wang

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 本研究提出面向放射学的开源VL基础模型RadFound,通过增强视觉编码器与跨模态学习设计,在自主构建的真实世界基准RadVLBench上显著优于其他VL模型,具备专家级放射学多模态处理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21956 2026-08-17 cs.CV 版本更新 77%

Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation

全局-局部双感知用于高分辨率文本密集图像翻译的MLLMs

Junxin Lu, Tengfei Song, Zhanglin Wu, Pengfei Li, Xiaowei Liang, Hui Yang, Kun Chen, Ning Xie, Yunfei Lu, Jing Zhao, Shiliang Sun, Daimeng Wei

机构 * School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院) Labs, Huawei Technologies Co., LTD(华为技术有限公司2012实验室) School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院)

专题命中 图文多模态 :multimodal(abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出GLoTran框架,通过全局-局部双感知方法提升高分辨率文本密集图像翻译的准确性和完整性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08418 2026-08-11 cs.CV 新提交 77%

Learning Deep Modality-Shared Self-Expressiveness for Image Clustering with Textual Information

用于结合文本信息的图像聚类的深度模态共享自表达学习

Xianghan Meng, Wei He, Zhiyuan Huang, Chun-Guang Li

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 提出DeepMORSE方法,通过模态共享自表达模型学习结合文本信息的图像聚类,在6个基准数据集上聚类性能提升超3%,且所学表示可迁移至下游任务。

详情

展开后加载摘要…

URL PDF HTML 收藏