arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4979 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4979 篇

2312.14385 2024-05-07 cs.DC cs.LG cs.MM 79%

Generative AI Beyond LLMs: System Implications of Multi-Modal Generation

Alicia Golden, Samuel Hsia, Fei Sun, Bilge Acun, Basil Hosmer, Yejin Lee, Zachary DeVito, Jeff Johnson, Gu-Yeon Wei, David Brooks, Carole-Jean Wu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.MM

Comments Published at 2024 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19444 2024-05-03 cs.CV 79%

AnomalyXFusion: Multi-modal Anomaly Synthesis with Diffusion

Jie Hu, Yawen Huang, Yilin Lu, Guoyang Xie, Guannan Jiang, Yefeng Zheng, Zhichao Lu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19175 2024-05-01 cs.CL 79%

Game-MUG: Multimodal Oriented Game Situation Understanding and Commentary Generation Dataset

Zhihao Zhang, Feiqi Cao, Yingbin Mo, Yiran Zhang, Josiah Poon, Caren Han

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14604 2024-04-29 cs.CL 79%

Describe-then-Reason: Improving Multimodal Mathematical Reasoning through Visual Comprehension Training

Mengzhao Jia, Zhihan Zhang, Wenhao Yu, Fangkai Jiao, Meng Jiang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16678 2024-04-26 cs.CV 79%

Multimodal Semantic-Aware Automatic Colorization with Diffusion Prior

Han Wang, Xinning Chai, Yiwen Wang, Yuhong Zhang, Rong Xie, Li Song

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15014 2024-04-24 cs.CV 79%

OccGen: Generative Multi-modal 3D Occupancy Prediction for Autonomous Driving

Guoqing Wang, Zhongdao Wang, Pin Tang, Jilai Zheng, Xiangxuan Ren, Bailan Feng, Chao Ma

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14567 2024-04-24 cs.CL 79%

WangLab at MEDIQA-M3G 2024: Multimodal Medical Answer Generation using Large Language Models

Ronald Xie, Steven Palayew, Augustin Toma, Gary Bader, Bo Wang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14367 2024-04-23 q-bio.QM cs.CL cs.LG 79%

Prot2Text: Multimodal Protein's Function Generation with GNNs and Transformers

Hadi Abdine, Michail Chatzianastasis, Costas Bouyioukos, Michalis Vazirgiannis

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 38(10), 10757-10765 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01284 2024-04-02 cs.CV 79%

Large Motion Model for Unified Multi-Modal Motion Generation

Mingyuan Zhang, Daisheng Jin, Chenyang Gu, Fangzhou Hong, Zhongang Cai, Jingfang Huang, Chongzhi Zhang, Xinying Guo, Lei Yang, Ying He, Ziwei Liu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Homepage: https://mingyuan-zhang.github.io/projects/LMM.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14828 2024-03-26 cs.CV 79%

Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing

Alberto Baldrati, Davide Morelli, Marcella Cornia, Marco Bertini, Rita Cucchiara

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11700 2024-03-25 cs.MM 79%

Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing

Juan Zhang, Jiahao Chen, Cheng Wang, Zhiwang Yu, Tangquan Qi, Can Liu, Di Wu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14141 2024-03-22 cs.CV 79%

Empowering Segmentation Ability to Multi-modal Large Language Models

Yuqi Yang, Peng-Tao Jiang, Jing Wang, Hao Zhang, Kai Zhao, Jinwei Chen, Bo Li

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16274 2024-03-22 cs.CV 79%

Towards Flexible, Scalable, and Adaptive Multi-Modal Conditioned Face Synthesis

Jingjing Ren, Cheng Xu, Haoyu Chen, Xinran Qin, Lei Zhu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08460 2024-03-20 cs.CV cs.RO 79%

Towards Dense and Accurate Radar Perception Via Efficient Cross-Modal Diffusion Model

Ruibin Zhang, Donglai Xue, Yuhan Wang, Ruixu Geng, Fei Gao

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

Comments 8 pages, 6 figures, submitted to RA-L

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.17336 2024-03-13 cs.CV cs.RO 79%

Robust 3D Object Detection from LiDAR-Radar Point Clouds via Cross-Modal Feature Augmentation

Jianning Deng, Gabriel Chan, Hantao Zhong, Chris Xiaoxuan Lu

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to ICRA 2024. 8 pages, 4 figures. Equal contribution for Gabriel Chan and Hantao Zhong, listed randomly

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06470 2024-03-12 cs.CV 79%

3D-aware Image Generation and Editing with Multi-modal Conditions

Bo Li, Yi-ke Li, Zhi-fen He, Bin Liu, Yun-Kun Lai

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04290 2024-03-08 eess.IV cs.CV cs.LG 79%

MedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual Invariant

Chenlu Zhan, Yu Lin, Gaoang Wang, Hongwei Wang, Jian Wu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04014 2024-03-08 cs.HC cs.AI 79%

PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement

Zhijie Wang, Yuheng Huang, Da Song, Lei Ma, Tianyi Zhang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments To appear in the 2024 CHI Conference on Human Factors in Computing Systems (CHI '24), May 11--16, 2024, Honolulu, HI, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10855 2024-02-19 cs.CV 79%

Control Color: Multimodal Diffusion-based Interactive Image Colorization

Zhexin Liang, Zhaochen Li, Shangchen Zhou, Chongyi Li, Chen Change Loy

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Project Page: https://zhexinliang.github.io/Control_Color/; Demo Video: https://youtu.be/tSCwA-srl8Q

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09036 2024-02-15 cs.CV 79%

Can Text-to-image Model Assist Multi-modal Learning for Visual Recognition with Visual Modality Missing?

Tiantian Feng, Daniel Yang, Digbalay Bose, Shrikanth Narayanan

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05803 2024-02-09 cs.CV cs.GR 79%

AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning

Wamiq Reyaz Para, Abdelrahman Eldesokey, Zhenyu Li, Pradyumna Reddy, Jiankang Deng, Peter Wonka

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06911 2024-02-07 cs.LG cs.CL q-bio.BM 79%

GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text

Pengfei Liu, Yiming Ren, Jun Tao, Zhixiang Ren

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments The article has been accepted by Computers in Biology and Medicine, with 14 pages and 4 figures

Journal ref Computers in Biology and Medicine, 108073, 2024, ISSN 0010-4825

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00375 2024-02-02 eess.IV cs.CV 79%

Disentangled Multimodal Brain MR Image Translation via Transformer-based Modality Infuser

Jihoon Cho, Xiaofeng Liu, Fangxu Xing, Jinsong Ouyang, Georges El Fakhri, Jinah Park, Jonghye Woo

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17664 2024-02-01 cs.CV cs.GR 79%

Image Anything: Towards Reasoning-coherent and Training-free Multi-modal Image Generation

Yuanhuiyi Lyu, Xu Zheng, Lin Wang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01827 2024-01-04 cs.CV 79%

Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions

David Junhao Zhang, Dongxu Li, Hung Le, Mike Zheng Shou, Caiming Xiong, Doyen Sahoo

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments project page: https://showlab.github.io/Moonshot/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17611 2024-01-01 cs.CV 79%

P2M2-Net: Part-Aware Prompt-Guided Multimodal Point Cloud Completion

Linlian Jiang, Pan Chen, Ye Wang, Tieru Wu, Rui Ma

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Best Poster Award of CAD/Graphics 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15900 2023-12-27 cs.CV 79%

Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional Control

Zunnan Xu, Yachao Zhang, Sicheng Yang, Ronghui Li, Xiu Li

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments AAAI-2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04445 2023-12-19 cs.LG cs.CV 79%

Multi-modal Latent Diffusion

Mustapha Bounoua, Giulio Franzese, Pietro Michiardi

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10512 2023-12-19 cs.CL cs.HC 79%

IMAD: IMage-Augmented multi-modal Dialogue

Viktor Moskvoretskii, Anton Frolov, Denis Kuznetsov

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Main part contains 6 pages, 4 figures. It was accepted on AINL. We wait the publication and DOI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08762 2023-12-15 cs.AI 79%

Multi-modal Latent Space Learning for Chain-of-Thought Reasoning in Language Models

Liqi He, Zuchao Li, Xiantao Cai, Ping Wang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏