arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4975 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4975 篇

2410.03860 2025-06-03 cs.CV 79%

MDMP: Multi-modal Diffusion for supervised Motion Predictions with uncertainty

Leo Bringer, Joey Wilson, Kira Barton, Maani Ghaffari

机构 * University of Michigan(密歇根大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2025 - HuMoGen. Minor revisions made based on reviewer feedback

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09724 2025-06-03 cs.CV 79%

MFCLIP: Multi-modal Fine-grained CLIP for Generalizable Diffusion Face Forgery Detection

Yaning Zhang, Tianyi Wang, Zitong Yu, Zan Gao, Linlin Shen, Shengyong Chen

机构 * Faculty of Computer Science and Technology, Qilu University of Technology (Shandong Academy of Sciences)(计算机科学与技术学院,齐鲁工业大学(山东科学院)) School of Computing, National University of Singapore(computing 学院,新加坡国立大学) School of Computing and Information Technology, Great Bay University(computing 与信息学院,大湾大学) National Engineering Laboratory for Big Data System Computing Technology, Shenzhen University(大数据系统计算技术国家工程实验室,深圳大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Information Forensics and Security 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24260 2025-06-02 cs.AI 79%

Generative AI for Urban Design: A Stepwise Approach Integrating Human Expertise with Multimodal Diffusion Models

Mingyi He, Yuebing Liang, Shenhao Wang, Yunhan Zheng, Qingyi Wang, Dingyi Zhuang, Li Tian, Jinhua Zhao

机构 * Department of Civil and Environmental Engineering, University of California, Berkeley(加州大学伯克利分校土木与环境工程系) The Singapore-MIT Alliance for Research and Technology(新加坡-麻省理工联盟研究技术中心) Department of Urban and Regional Planning, University of Florida(佛罗里达大学城市与区域规划系) Department of Civil and Environmental Engineering, Massachusetts Institute of Technology(麻省理工学院土木与环境工程系) Department of Urban Planning, Tsinghua University(清华大学城市规划系) Department of Urban Studies and Planning, Massachusetts Institute of Technology(麻省理工学院城市研究与规划系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22948 2025-05-30 cs.AI 79%

Foundation Molecular Grammar: Multi-Modal Foundation Models Induce Interpretable Molecular Graph Languages

Michael Sun, Weize Yuan, Gang Liu, Wojciech Matusik, Jie Chen

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) MIT Chemistry(麻省理工学院化学系) MIT-IBM Watson AI Lab, IBM Research(麻省理工-IBM Watson人工智能实验室,IBM研究院) University of Notre Dame(诺埃伯大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.09117 2025-05-30 cs.CV 79%

Cross-Modal Causal Intervention for Medical Report Generation

Weixing Chen, Yang Liu, Ce Wang, Jiarui Zhu, Guanbin Li, Cheng-Lin Liu, Liang Lin

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室) School of Science, Sun Yat-sen University(中山大学理学院) Hong Kong Polytechnic University(香港理工大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by IEEE TIP 2025, 16 pages, 11 figures, 7 tables

Journal ref IEEE Transactions on Image Processing 34 (2025) 2970-2985

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16990 2025-05-27 cs.CV 79%

Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding

Runpeng Yu, Xinyin Ma, Xinchao Wang

机构 * National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20115 2025-05-27 cs.SE cs.AI 79%

AutoP2C: An LLM-Based Agent Framework for Code Repository Generation from Multimodal Content in Academic Papers

Zijie Lin, Yiqing Shen, Qilin Cai, He Sun, Jinrui Zhou, Mingjun Xiao

机构 * University of Science and Technology of China(中国科学技术大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18341 2025-05-27 cs.RO cs.AI 79%

CrashAgent: Crash Scenario Generation via Multi-modal Reasoning

Miao Li, Wenhao Ding, Haohong Lin, Yiqi Lyu, Yihang Yao, Yuyou Zhang, Ding Zhao

机构 * Carnegie Mellon University(卡内基梅隆大学) NVIDIA Research(NVIDIA研究) Northwestern University(西北大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16602 2025-05-23 cs.CV 79%

MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation

Bohan Zhou, Yi Zhan, Zhongbin Zhang, Zongqing Lu

机构 * School of Computer Science, Peking University(北京大学计算机学院) Department of Automation, Tsinghua University(清华大学自动化系) BeingBeyond

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12606 2025-05-20 cs.CV 79%

Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking

Shiyu Xuan, Zechao Li, Jinhui Tang

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(计算机科学与工程学院,南京理工大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13175 2025-05-20 cs.LG cs.AI physics.ao-ph 79%

TCP-Diffusion: A Multi-modal Diffusion Model for Global Tropical Cyclone Precipitation Forecasting with Change Awareness

Cheng Huang, Pan Mu, Cong Bai, Peter AG Watson

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments Camera-ready version. This paper has been accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08803 2025-05-15 cs.LG cs.AI 79%

Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models

Zizhao Hu, Mohammad Rostami, Jesse Thomason

机构 * University of Southern California(南加州大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06303 2025-05-13 cs.LG cs.AI 79%

Collaborative Multi-LoRA Experts with Achievement-based Multi-Tasks Loss for Unified Multimodal Information Extraction

Li Yuan, Yi Cai, Xudong Shen, Qing Li, Qingbao Huang, Zikun Deng, Tao Wang

机构 * School of Software Engineering, South China University of Technology, Guangzhou, China(华南理工大学软件学院) Key Laboratory of Big Data and Intelligent Robot (SCUT), MOE of China(大数据与智能机器人重点实验室) Department of Computing, The Hong Kong Polytechnic University, Hong Kong, China(香港理工大学计算机系) School of Electrical Engineering, Guangxi University, Nanning, China(广西大学电气工程学院) Department of Biostatistics & Health Informatics, King’s College London, London, United Kingdom(伦敦国王学院生物统计与健康信息学系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04650 2025-05-09 cs.GR cs.AI cs.IR cs.LG 79%

Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models

Kapil Wanaskar, Gaytri Jena, Magdalini Eirinaki

机构 * Computer Engineering Dept. San José State University(计算机工程系圣何塞州立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04996 2025-05-09 cs.CL 79%

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Weixin Liang, Lili Yu, Liang Luo, Srinivasan Iyer, Ning Dong, Chunting Zhou, Gargi Ghosh, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, Xi Victoria Lin

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted to TMLR 2025; 48 pages

Journal ref Transactions on Machine Learning Research (2025), ISSN: 2835-8856

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21718 2025-05-08 cs.CV 79%

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction

Shiying Li, Xingqun Qi, Bingkun Yang, Chen Weile, Zezhao Tian, Muyi Sun, Qifeng Liu, Man Zhang, Zhenan Sun

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Hong Kong University of Science and Technology(香港科技大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19341 2025-04-29 cs.RO cs.AI 79%

PolyTouch: A Robust Multi-Modal Tactile Sensor for Contact-rich Manipulation Using Tactile-Diffusion Policies

Jialiang Zhao, Naveen Kuppuswamy, Siyuan Feng, Benjamin Burchfiel, Edward Adelson

机构 * MIT CSAIL

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments Nominated for the best paper award at ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14320 2025-04-23 cs.HC cs.AI 79%

Expanding the Generative AI Design Space through Structured Prompting and Multimodal Interfaces

Nimisha Karnatak, Adrien Baranes, Rob Marchant, Huinan Zeng, Tríona Butler, Kristen Olson

机构 * University of Oxford(牛津大学) Google DeepMind(谷歌DeepMind) King’s College London(伦敦大学学院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at CHI'25 Workshop on Designing and Developing User Interfaces with AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21979 2025-04-23 cs.CV 79%

Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Size Wu, Wenwei Zhang, Lumin Xu, Sheng Jin, Zhonghua Wu, Qingyi Tao, Wentao Liu, Wei Li, Chen Change Loy

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) Shanghai AI Laboratory Research(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) SenseTime Research and Tetras.AI(SenseTime研究部和Tetras.AI) SenseTime Research(商汤科技研究院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03173 2025-04-23 cs.CV 79%

MObI: Multimodal Object Inpainting Using Diffusion Models

Alexandru Buburuzan, Anuj Sharma, John Redford, Puneet K. Dokania, Romain Mueller

机构 * FiveAI The University of Manchester(曼彻斯特大学) University of Oxford(牛津大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages; Project page at https://alexbubu.com/mobi

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14666 2025-04-22 cs.CV 79%

Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

Kaihang Pan, Wang Lin, Zhongqi Yue, Tenglong Ao, Liyu Jia, Wei Zhao, Juncheng Li, Siliang Tang, Hanwang Zhang

机构 * Zhejiang University(浙江大学) Nanyang Technological University(新加坡国立大学) Peking University(北京大学) Huawei Singapore Research Center(华为新加坡研究中心)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04161 2025-04-22 cs.CV 79%

Overcoming False Illusions in Real-World Face Restoration with Multi-Modal Guided Diffusion Model

Keda Tao, Jinjin Gu, Yulun Zhang, Xiucheng Wang, Nan Cheng

机构 * Xidian University(西安电子科技大学) The University of Sydney(悉尼大学) Shanghai Jiao Tong University(上海交通大学) State Key Laboratory of Integrated Services Networks(集成服务网络国家重点实验室)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 23 Pages, 28 Figures, ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13631 2025-04-21 cs.AI 79%

Multi-modal Knowledge Graph Generation with Semantics-enriched Prompts

Yajing Xu, Zhiqiang Liu, Jiaoyan Chen, Mingchen Tu, Zhuo Chen, Jeff Z. Pan, Yichi Zhang, Yushan Zhu, Wen Zhang, Huajun Chen

机构 * Zhejiang University(浙江大学) Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系) School of Informatics, The University of Edinburgh(爱丁堡大学信息学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted by IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08857 2025-04-21 cs.CV 79%

DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation

Minbin Huang, Yanxin Long, Xinchi Deng, Ruihang Chu, Jiangfeng Xiong, Xiaodan Liang, Hong Cheng, Qinglin Lu, Wei Liu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Project page: https://hunyuan-dialoggen.github.io/. Accepted to NAACL2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12844 2025-04-18 cs.CV 79%

High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion

Libo Zhang, Yongsheng Yu, Jiali Yao, Heng Fan

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to IJCV. arXiv admin note: text overlap with arXiv:2208.11850

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04110 2025-04-17 cs.HC cs.AI 79%

InterChat: Enhancing Generative Visual Analytics using Multimodal Interactions

Juntong Chen, Jiang Wu, Jiajing Guo, Vikram Mohanty, Xueming Li, Jorge Piazentin Ono, Wenbin He, Liu Ren, Dongyu Liu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments This work is accepted by the 27th Eurographics Conference on Visualization (EuroVis 2025). The paper contains 12 pages and 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08358 2025-04-14 cs.CV 79%

LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs

Jiarui Wang, Huiyu Duan, Yu Zhao, Juntong Wang, Guangtao Zhai, Xiongkuo Min

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08111 2025-04-14 cs.CV 79%

POEM: Precise Object-level Editing via MLLM control

Marco Schouten, Mehmet Onurcan Kaya, Serge Belongie, Dim P. Papadopoulos

专题命中 多模态生成 :MLLM(title);multimodal(abstract);分类 cs.CV

Comments Accepted to SCIA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13947 2025-04-14 cs.CV 79%

Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation

Sayak Nag, Udita Ghosh, Calvin-Khang Ta, Sarosij Bose, Jiachen Li, Amit K Roy Chowdhury

专题命中 多模态生成 :MLLM(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05821 2025-04-14 cs.CV 79%

F-LMM: Grounding Frozen Large Multimodal Models

Size Wu, Sheng Jin, Wenwei Zhang, Lumin Xu, Wentao Liu, Wei Li, Chen Change Loy

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Project Page: https://github.com/wusize/F-LMM

详情

展开后加载摘要…

URL PDF HTML 收藏