arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4975 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4975 篇

2407.02994 2025-07-18 cs.DB cs.AI cs.LG 79%

MedPix 2.0: A Comprehensive Multimodal Biomedical Data set for Advanced AI Applications with Retrieval Augmented Generation and Knowledge Graphs

Irene Siragusa, Salvatore Contino, Massimo La Ciura, Rosario Alicata, Roberto Pirrone

机构 * Department of Engineering, University of Palermo(1 工程学院,巴勒莫大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Journal ref Data Sci. Eng. (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06662 2025-07-10 cs.CV cs.RO 79%

MK-Pose: Category-Level Object Pose Estimation via Multimodal-Based Keypoint Learning

Yifan Yang, Peili Song, Enfan Lan, Dong Liu, Jingtai Liu

机构 * Institute of Robotics and Automatic Information System(机器人与自动信息系统研究所) Tianjin Key Laboratory of Intelligent Robotics(天津智能机器人重点实验室) TBI center(TBI中心) Nankai University(南开大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06571 2025-07-10 cs.CL 79%

Enhancing Food-Domain Question Answering with a Multimodal Knowledge Graph: Hybrid QA Generation and Diversity Analysis

Srihari K B, Pushpak Bhattacharyya

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05790 2025-07-09 cs.CV 79%

TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model

Yujie Hu, Xuanyu Zhang, Weiqi Li, Jian Zhang

机构 * School of Electronic and Computer Engineering, Peking University, China(北京大学电子与计算机工程学院) Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology, Shenzhen Graduate School, Peking University(北京大学深圳研究生院)

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments 6 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02995 2025-07-09 cs.CV cs.CR 79%

FreqCross: A Multi-Modal Frequency-Spatial Fusion Network for Robust Detection of Stable Diffusion 3.5 Generated Images

Guang Yang

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04000 2025-07-08 cs.IR cs.AI 79%

Leveraging Multimodal Data and Side Users for Diffusion Cross-Domain Recommendation

Fan Zhang, Jinpeng Chen, Huan Li, Senzhang Wang, Yuan Cao, Kaimin Wei, JianXiang He, Feifei Kou, Jinqing Wang

机构 * School of Computer Science (National Pilot Software Engineering School), Beijing University of Posts and Telecommunications(北京邮电大学计算机学院(国家级试点软件工程学院)) The State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江大学区块链与数据安全国家重点实验室) Central South University(中南大学) National University of Singapore(新加坡国立大学) Jinan University(暨南大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14453 2025-07-08 cs.CV cs.GR cs.LG 79%

Multimodal Latent Diffusion Model for Complex Sewing Pattern Generation

Shengqi Liu, Yuhao Cheng, Zhuo Chen, Xingyu Ren, Wenhan Zhu, Lincheng Li, Mengxiao Bi, Xiaokang Yang, Yichao Yan

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University, China(人工智能领域关键实验室、人工智能研究院、上海交通大学)

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Our project page: https://shengqiliu1.github.io/SewingLDM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03326 2025-07-08 cs.CV 79%

Mirror in the Model: Ad Banner Image Generation via Reflective Multi-LLM and Multi-modal Agents

Zhao Wang, Bowen Chen, Yotaro Shimose, Sota Moriyama, Heng Wang, Shingo Takamatsu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22374 2025-06-30 cs.LG cs.AI 79%

Sheaf-Based Decentralized Multimodal Learning for Next-Generation Wireless Communication Systems

Abdulmomen Ghalkha, Zhuojun Tian, Chaouki Ben Issaid, Mehdi Bennis

机构 * Center for Wireless Communications, University of Oulu(无线通信中心,奥卢大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21832 2025-06-30 cs.CV 79%

TaleForge: Interactive Multimodal System for Personalized Story Creation

Minh-Loi Nguyen, Quang-Khai Le, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * 1 University of Science, VNU-HCM, Ho Chi Minh City, Vietnam 2 Vietnam National University, Ho Chi Minh City, Vietnam 3 Department of Computer Science \ of Dayton Ohio, United States

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20155 2025-06-26 cs.CV 79%

Towards Efficient Exemplar Based Image Editing with Multimodal VLMs

Avadhoot Jadhav, Ashutosh Srivastava, Abhinav Java, Silky Singh, Tarun Ram Menta, Surgan Jandial, Balaji Krishnamurthy

机构 * Indian Institute of Technology, Bombay(印度理工学院班加罗尔分校) Indian Institute of Technology, Roorkee(印度理工学院罗尔基分校) Microsoft Research(微软研究院) Stanford University(斯坦福大学) Adobe MDSR(Adobe MDSR实验室) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ECCV 2024 (AI4VA Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00659 2025-06-26 cs.RO cs.AI 79%

Multimodal Coherent Explanation Generation of Robot Failures

Pradip Pramanick, Silvia Rossi

机构 * Interdepartmental Center for Advances in Robotic Surgery - ICAROS, University of Naples Federico II(跨部门先进机器人手术中心 - ICAROS,那不勒斯费德里科二世大学) Department of Electrical Engineering and Information Technologies - DIETI, University of Naples Federico II(电气工程与信息科技系 - DIETI,那不勒斯费德里科二世大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Journal ref 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11299 2025-06-25 cs.CV 79%

MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching

Yepeng Liu, Zhichao Sun, Baosheng Yu, Yitian Zhao, Bo Du, Yongchao Xu, Jun Cheng

机构 * National Engineering Research Center for Multimedia Software(多媒体软件国家工程研究中心) Institute of Artificial Intelligence(人工智能研究院) School of Computer Science(计算机科学学院) Hubei Key Laboratory of Multimedia and Network Communication Engineering(湖北省多媒体与网络通信工程重点实验室) Lee Kong Chian School of Medicine(李光耀医学院) Nanyang Technological University(南洋理工大学) Ningbo Institute of Materials Technology and Engineering(宁波材料技术与工程研究所) Chinese Academy of Sciences(中国科学院) Institute for Infocomm Research (I 2 R)(信息与通信研究所以(I 2 R)) Agency for Science, Technology and Research (A*STAR)(科技研究局(A*STAR))

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accept by IEEE TIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17645 2025-06-24 cs.CV 79%

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning

Shih-Wen Liu, Hsuan-Yu Fan, Wei-Ta Chu, Fu-En Yang, Yu-Chiang Frank Wang

机构 * National Cheng Kung University, Taiwan(国立成功大学) NVIDIA Research(NVIDIA研究)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to MIDL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05298 2025-06-23 cs.CL 79%

Coreference as an indicator of context scope in multimodal narrative

Nikolai Ilinykh, Shalom Lappin, Asad Sayeed, Sharid Loáiciga

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments 19 pages, 4 tables. Accepted to GEM2 Workshop: Generation, Evaluation & Metrics at ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15372 2025-06-19 cs.CL 79%

COSMMIC: Comment-Sensitive Multimodal Multilingual Indian Corpus for Summarization and Headline Generation

Raghvendra Kumar, S. A. Mohammed Salman, Aryan Sahu, Tridib Nandi, Pragathi Y. P., Sriparna Saha, Jose G. Moreno

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Patna, India(印度理工学院帕纳分校计算机科学与工程系) Department of Metallurgical and Materials Engineering, National Institute of Technology Tiruchirappalli, India(印度理工学院 Tiruchirappalli 金属与材料工程系) Department of Computer Science and Information Systems, BITS Pilani – Goa Campus, India(比斯·帕尼学院 Goa 分校计算机科学与信息系统系) Department of Computer Science and Engineering, Indian Institute of Information Technology Vadodara, India(印度信息科技学院瓦达拉分校计算机科学与工程系) Department of Computer Science and Engineering, B.M.S. College of Engineering, Bangalore, India(班加罗尔 B.M.S. 工程学院计算机科学与工程系) Université de Toulouse, IRIT UMR 5505 CNRS, France(图卢兹大学 IRIT UMR 5505 CNRS 实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments ACL 2025 MAINs

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13667 2025-06-17 eess.IV cs.CV 79%

MultiViT2: A Data-augmented Multimodal Neuroimaging Prediction Framework via Latent Diffusion Model

Bi Yuda, Jia Sihan, Gao Yutong, Abrol Anees, Fu Zening, Calhoun Vince

机构 * Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS)(跨机构神经影像与数据科学转化研究中心(TReNDS))

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11178 2025-06-16 cs.CV cs.LG cs.NE 79%

BrainMAP: Multimodal Graph Learning For Efficient Brain Disease Localization

Nguyen Linh Dan Le, Jing Ren, Ciyuan Peng, Chengyao Xie, Bowen Li, Feng Xia

机构 * School of Computing Technologies, RMIT University(计算技术学院,拉筹纳斯大学) Institute of Innovation, Science and Sustainability, Federation University Australia(创新、科学与可持续性研究所,联邦大学澳大利亚)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 6 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09182 2025-06-16 eess.IV cs.CV 79%

seg2med: a bridge from artificial anatomy to multimodal medical images

Zeyu Yang, Zhilin Chen, Yipeng Sun, Anika Strittmatter, Anish Raj, Ahmad Allababidi, Johann S. Rink, Frank G. Zöllner

机构 * Computer Assisted Clinical Medicine, Medical Faculty Mannheim, Heidelberg University(计算机辅助临床医学,曼海姆医学院,海德堡大学) Pattern Recognition Lab, Friedrich-Alexander-University Erlangen-Nuremberg(模式识别实验室,埃尔兰根-纽伦堡弗里德里希-亚历山大大学) Department of Radiology and Nuclear Medicine, University Medical Center Mannheim(放射学与核医学系,曼海姆大学医学中心) Mannheim Institute for Intelligent Systems in Medicine, Medical Faculty Mannheim, Heidelberg University(曼海姆智能医学研究所,曼海姆医学院,海德堡大学) Optical Bioimaging Laboratory, Department of Biomedical Engineering, College of Design and Engineering, National University of Singapore(光学生物成像实验室,生物医学工程系,设计与工程学院,新加坡国立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 17 pages, 10 figures Web demo available at https://huggingface.co/spaces/Zeyu0601/frankenstein

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06733 2025-06-12 cs.CV 79%

RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation

Ruoxuan Zhang, Jidong Gao, Bin Wen, Hongxia Xie, Chenming Zhang, Hong-Han Shuai, Wen-Huang Cheng

机构 * Jilin University(吉林大学) Guangdong University of Technology(广东工业大学) National Chiao Tung University(交通大学) National Taiwan University(国立台湾大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments This is an extended version of arXiv:2503.05228

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01343 2025-06-11 cs.AI 79%

BalancEdit: Dynamically Balancing the Generality-Locality Trade-off in Multi-modal Model Editing

Dongliang Guo, Mengxuan Hu, Zihan Guan, Thomas Hartvigsen, Sheng Li

机构 * University of Virginia(弗吉尼亚大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Journal ref Proceedings of the 42 nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05153 2025-06-10 cs.CV 79%

Hummingbird: High Fidelity Image Generation via Multimodal Context Alignment

Minh-Quan Le, Gaurav Mittal, Tianjian Meng, A S M Iftekhar, Vishwas Suryanarayanan, Barun Patra, Dimitris Samaras, Mei Chen

机构 * Microsoft(微软公司) Stony Brook University(史蒂文尼森布鲁克大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to ICLR 2025. Project page with code release: https://roar-ai.github.io/hummingbird

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16656 2025-06-09 cs.CV 79%

Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

Peiyu Wang, Yichen Wei, Yi Peng, Xiaokun Wang, Weijie Qiu, Wei Shen, Tianyidan Xie, Jiangbo Pei, Jianhao Zhang, Yunzhuo Hao, Xuchen Song, Yang Liu, Yahui Zhou

机构 * Skywork AI Kunlun Inc.(Kunlun公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10798 2025-06-05 cs.CV 79%

MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling

Jian Yang, Dacheng Yin, Yizhou Zhou, Fengyun Rao, Wei Zhai, Yang Cao, Zheng-Jun Zha

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发式智能感知与认知大学科学与技术研究院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03007 2025-06-04 cs.CV 79%

DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models

Jiarui Wang, Huiyu Duan, Juntong Wang, Ziheng Jia, Woo Yi Yang, Xiaorong Zhu, Yu Zhao, Jiaying Qian, Yuke Xing, Guangtao Zhai, Xiongkuo Min

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02467 2025-06-04 eess.IV cs.CV 79%

Multi-modal brain MRI synthesis based on SwinUNETR

Haowen Pang, Weiyan Guo, Chuyang Ye

机构 * School of Integrated Circuits and Electronics, Beijing Institute of Technology, Beijing, China(集成电路与电子学院,北京理工大学,北京,中国)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19901 2025-06-04 cs.CV 79%

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM

Peng Liu, Xiaoming Ren, Fengkai Liu, Qingsong Xie, Quanlong Zheng, Yanhao Zhang, Haonan Lu, Yujiu Yang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02433 2025-06-04 cs.CV 79%

OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style Mimicking

Zhongjian Wang, Peng Zhang, Jinwei Qi, Guangyuan Wang, Chaonan Ji, Sheng Xu, Bang Zhang, Liefeng Bo

机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)

专题命中 多模态生成 :multimodal(title);cross-modal(abstract);分类 cs.CV

Comments Project Page https://humanaigc.github.io/omnitalker

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01853 2025-06-03 cs.CV 79%

ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

Junliang Ye, Zhengyi Wang, Ruowen Zhao, Shenghao Xie, Jun Zhu

机构 * Tsinghua University(清华大学) Peking University(北京大学) ShengShu(盛舒)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Project page: https://github.com/JAMESYJL/ShapeLLM-Omni

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02640 2025-06-03 cs.MM 79%

RoSMM: A Robust and Secure Multi-Modal Watermarking Framework for Diffusion Models

ZhongLi Fang, Yu Xie, Ping Chen

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏