arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4979 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4979 篇

2407.14936 2024-07-23 cs.MM 79%

EidetiCom: A Cross-modal Brain-Computer Semantic Communication Paradigm for Decoding Visual Perception

Linfeng Zheng, Peilin Chen, Shiqi Wang

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11213 2024-07-17 cs.CV 79%

OpenPSG: Open-set Panoptic Scene Graph Generation via Large Multimodal Models

Zijian Zhou, Zheng Zhu, Holger Caesar, Miaojing Shi

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08473 2024-07-12 cs.AR cs.AI 79%

Natural language is not enough: Benchmarking multi-modal generative AI for Verilog generation

Kaiyan Chang, Zhirong Chen, Yunhao Zhou, Wenlong Zhu, kun wang, Haobo Xu, Cangyuan Li, Mengdi Wang, Shengwen Liang, Huawei Li, Yinhe Han, Ying Wang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted by ICCAD 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05808 2024-07-10 cs.CV eess.IV 79%

Adaptive Multi-modal Fusion of Spatially Variant Kernel Refinement with Diffusion Model for Blind Image Super-Resolution

Junxiong Lin, Yan Wang, Zeng Tao, Boyang Wang, Qing Zhao, Haorang Wang, Xuan Tong, Xinji Mai, Yuxuan Lin, Wei Song, Jiawen Yu, Shaoqi Yan, Wenqiang Zhang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05340 2024-07-10 cs.CV eess.IV 79%

Unified Multi-Modal Image Synthesis for Missing Modality Imputation

Yue Zhang, Chengtao Peng, Qiuli Wang, Dan Song, Kaiyan Li, S. Kevin Zhou

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments IEEE TMI accepted final version

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04736 2024-07-09 eess.SP cs.AI cs.LG 79%

SCDM: Unified Representation Learning for EEG-to-fNIRS Cross-Modal Generation in MI-BCIs

Yisheng Li, Shuqiang Wang

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.AI

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00621 2024-07-04 cs.IR cs.MM 79%

Multimodal Pretraining, Adaptation, and Generation for Recommendation: A Survey

Qijiong Liu, Jieming Zhu, Yanting Yang, Quanyu Dai, Zhaocheng Du, Xiao-Ming Wu, Zhou Zhao, Rui Zhang, Zhenhua Dong

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.MM

Comments Accepted by KDD 2024. See our tutorial materials at https://mmrec.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14109 2024-07-04 cs.AI 79%

Boosting the Power of Small Multimodal Reasoning Models to Match Larger Models with Self-Consistency Training

Cheng Tan, Jingxuan Wei, Zhangyang Gao, Linzhuang Sun, Siyuan Li, Ruifeng Guo, Bihui Yu, Stan Z. Li

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14676 2024-07-02 cs.CV cs.GR 79%

DreamPBR: Text-driven Generation of High-resolution SVBRDF with Multi-modal Guidance

Linxuan Xin, Zheng Zhang, Jinfu Wei, Wei Gao, Duan Gao

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 16 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00118 2024-07-02 cs.LG cs.AI 79%

From Efficient Multimodal Models to World Models: A Survey

Xinji Mai, Zeng Tao, Junxiong Lin, Haoran Wang, Yang Chang, Yanlan Kang, Yan Wang, Wenqiang Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18542 2024-06-28 cs.CV eess.SP 79%

Generative AI Empowered LiDAR Point Cloud Generation with Multimodal Transformer

Mohammad Farzanullah, Han Zhang, Akram Bin Sediq, Ali Afana, Melike Erol-Kantarci

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 6 pages, 4 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18247 2024-06-27 eess.IV cs.CV cs.LG 79%

Generative artificial intelligence in ophthalmology: multimodal retinal images for the diagnosis of Alzheimer's disease with convolutional neural networks

I. R. Slootweg, M. Thach, K. R. Curro-Tafili, F. D. Verbraak, F. H. Bouwman, Y. A. L. Pijnenburg, J. F. Boer, J. H. P. de Kwisthout, L. Bagheriye, P. J. González

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14555 2024-06-21 cs.CV 79%

A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Xincheng Shuai, Henghui Ding, Xingjun Ma, Rongcheng Tu, Yu-Gang Jiang, Dacheng Tao

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Project Page: https://github.com/xinchengshuai/Awesome-Image-Editing

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13150 2024-06-21 eess.IV cs.CV 79%

MCAD: Multi-modal Conditioned Adversarial Diffusion Model for High-Quality PET Image Reconstruction

Jiaqi Cui, Xinyi Zeng, Pinxian Zeng, Bo Liu, Xi Wu, Jiliu Zhou, Yan Wang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Early accepted by MICCAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10561 2024-06-18 cs.CL 79%

We Care: Multimodal Depression Detection and Knowledge Infused Mental Health Therapeutic Response Generation

Palash Moon, Pushpak Bhattacharyya

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.00822 2024-06-18 cs.IR cs.AI 79%

InteraRec: Screenshot Based Recommendations Using Multimodal Large Language Models

Saketh Reddy Karra, Theja Tulabandhula

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09162 2024-06-14 cs.CV 79%

EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal Prompts

Yucheng Han, Rui Wang, Chi Zhang, Juntao Hu, Pei Cheng, Bin Fu, Hanwang Zhang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments https://tencentqqgylab.github.io/EMMA

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09003 2024-06-14 cs.CV cs.LG 79%

Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation

Lincan Cai, Shuang Li, Wenxuan Ma, Jingxuan Kang, Binhui Xie, Zixun Sun, Chengwei Zhu

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06665 2024-06-14 cs.CV 79%

Deep Generative Data Assimilation in Multimodal Setting

Yongquan Qu, Juan Nathaniel, Shuolin Li, Pierre Gentine

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Best Student Paper Award @ CVPR2024 EarthVision

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05325 2024-06-11 eess.AS cs.SD 79%

LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance

Shihao Chen, Yu Gu, Jie Zhang, Na Li, Rilin Chen, Liping Chen, Lirong Dai

专题命中 多模态生成 :any-to-any(title,abstract);分类 eess.AS

Comments Accepted by Interspeech 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01334 2024-06-04 cs.CV 79%

HHMR: Holistic Hand Mesh Recovery by Enhancing the Multimodal Controllability of Graph Diffusion Models

Mengcheng Li, Hongwen Zhang, Yuxiang Zhang, Ruizhi Shao, Tao Yu, Yebin Liu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments accepted in CVPR2024, project page: https://dw1010.github.io/project/HHMR/HHMR.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16501 2024-05-28 cs.CV 79%

User-Friendly Customized Generation with Multi-Modal Prompts

Linhao Zhong, Yan Hong, Wentao Chen, Binglin Zhou, Yiyi Zhang, Jianfu Zhang, Liqing Zhang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 11 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15738 2024-05-27 cs.CV 79%

ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Chunjiang Ge, Sijie Cheng, Ziming Wang, Jiale Yuan, Yuan Gao, Jun Song, Shiji Song, Gao Huang, Bo Zheng

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12741 2024-05-27 cs.CV 79%

MuLan: Multimodal-LLM Agent for Progressive and Interactive Multi-Object Diffusion

Sen Li, Ruochen Wang, Cho-Jui Hsieh, Minhao Cheng, Tianyi Zhou

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Added the application to human-agent interaction; added discussion with concurrent work

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04834 2024-05-24 cs.CV 79%

FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation

Xuehai He, Jian Zheng, Jacob Zhiyuan Fang, Robinson Piramuthu, Mohit Bansal, Vicente Ordonez, Gunnar A Sigurdsson, Nanyun Peng, Xin Eric Wang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01655 2024-05-21 cs.CV 79%

FashionEngine: Interactive 3D Human Generation and Editing via Multimodal Controls

Tao Hu, Fangzhou Hong, Zhaoxi Chen, Ziwei Liu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Project Page: https://taohuumd.github.io/projects/FashionEngine

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02233 2024-05-13 cs.CV 79%

MedXChat: A Unified Multimodal Large Language Model Framework towards CXRs Understanding and Generation

Ling Yang, Zhanyu Wang, Zhenghao Chen, Xinyu Liang, Luping Zhou

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13286 2024-05-09 cs.CV 79%

Generative Multimodal Models are In-Context Learners

Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiying Yu, Zhengxiong Luo, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, Xinlong Wang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2024. Project page: https://baaivision.github.io/emu2

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04356 2024-05-08 cs.CV 79%

Diffusion-driven GAN Inversion for Multi-Modal Face Image Generation

Jihyun Kim, Changjae Oh, Hoseok Do, Soohyun Kim, Kwanghoon Sohn

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02911 2024-05-07 cs.CV 79%

Multimodal Sense-Informed Prediction of 3D Human Motions

Zhenyu Lou, Qiongjie Cui, Haofan Wang, Xu Tang, Hong Zhou

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏