arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4975 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4975 篇

2404.10763 2024-04-17 cs.AI cs.CL cs.CV 67%

LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?

Yuchi Wang, Shuhuai Ren, Rundong Gao, Linli Yao, Qingyan Guo, Kaikai An, Jianhong Bai, Xu Sun

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04552 2024-04-16 cs.CV cs.AI cs.LG cs.MM 67%

Generating Illustrated Instructions

Sachit Menon, Ishan Misra, Rohit Girdhar

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted to CVPR 2024. Project website: http://facebookresearch.github.io/IllustratedInstructions. Code reproduction: https://github.com/sachit-menon/generating-illustrated-instructions-reproduction

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07206 2024-04-11 cs.CV cs.AI cs.GR cs.LG cs.MM 67%

GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models

Zewei Zhang, Huan Liu, Jun Chen, Xiangyu Xu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.18814 2024-03-28 cs.CV cs.AI cs.CL 67%

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin Chen, Ruihang Chu, Shaoteng Liu, Jiaya Jia

专题命中 多模态生成 :any-to-any(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Code and models are available at https://github.com/dvlab-research/MiniGemini

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06952 2024-03-12 cs.CV cs.AI cs.CL cs.LG 67%

SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data

Jialu Li, Jaemin Cho, Yi-Lin Sung, Jaehong Yoon, Mohit Bansal

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments First two authors contributed equally; Project website: https://selma-t2i.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14623 2024-02-23 cs.RO cs.AI cs.CL cs.CV 67%

RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation

Junting Chen, Yao Mu, Qiaojun Yu, Tianming Wei, Silang Wu, Zhecheng Yuan, Zhixuan Liang, Chao Yang, Kaipeng Zhang, Wenqi Shao, Yu Qiao, Huazhe Xu, Mingyu Ding, Ping Luo

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 10 pages of main paper, 4 pages of appendix; 10 figures in main paper, 3 figures in appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.10134 2024-02-20 cs.MM cs.CL cs.CV 67%

Recipe Generation from Unsegmented Cooking Videos

Taichi Nishimura, Atsushi Hashimoto, Yoshitaka Ushiku, Hirotaka Kameko, Shinsuke Mori

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted at ACM TOMM; ACM Transactions on Multimedia Computing, Communications, and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19216 2024-02-16 cs.CL cs.AI cs.CV cs.LG 67%

Translation-Enhanced Multilingual Text-to-Image Generation

Yaoyiran Li, Ching-Yun Chang, Stephen Rawls, Ivan Vulić, Anna Korhonen

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL 2023 (Main)

Journal ref Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pages 9174-9193

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01044 2024-01-03 cs.SD cs.AI cs.CL eess.AS 67%

Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Jinlong Xue, Yayue Deng, Yingming Gao, Ya Li

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments Demo and implementation at https://auffusion.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19228 2023-12-25 cs.CL cs.AI cs.SD eess.AS 67%

Unsupervised Melody-to-Lyric Generation

Yufei Tian, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Gunnar Sigurdsson, Chenyang Tao, Wenbo Zhao, Yiwen Chen, Tagyoung Chung, Jing Huang, Nanyun Peng

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments ACL 2023. arXiv admin note: substantial text overlap with arXiv:2305.07760

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04749 2023-12-07 cs.CV cs.AI cs.LG cs.MM stat.ML 67%

Divide, Evaluate, and Refine: Evaluating and Improving Text-to-Image Alignment with Iterative VQA Feedback

Jaskirat Singh, Liang Zheng

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Journal ref Published at NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11097 2023-11-21 cs.CV cs.AI cs.CL cs.LG 67%

Radiology Report Generation Using Transformers Conditioned with Non-imaging Data

Nurbanu Aksoy, Nishant Ravikumar, Alejandro F Frangi

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16397 2023-11-06 cs.CV cs.AI cs.CL 67%

Are Diffusion Models Vision-And-Language Reasoners?

Benno Krojer, Elinor Poole-Dayan, Vikram Voleti, Christopher Pal, Siva Reddy

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12503 2023-09-12 cs.SD cs.AI cs.MM eess.AS eess.SP 67%

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, Mark D. Plumbley

专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI、cs.MM、eess.AS

Comments Accepted by ICML 2023. Demo and implementation at https://audioldm.github.io. Evaluation toolbox at https://github.com/haoheliu/audioldm_eval

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03549 2023-09-08 cs.CV cs.AI cs.MM 67%

Reuse and Diffuse: Iterative Denoising for Text-to-Video Generation

Jiaxi Gu, Shicong Wang, Haoyu Zhao, Tianyi Lu, Xing Zhang, Zuxuan Wu, Songcen Xu, Wei Zhang, Yu-Gang Jiang, Hang Xu

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07605 2023-08-16 cs.CV cs.AI cs.MM 67%

SGDiff: A Style Guided Diffusion Model for Fashion Synthesis

Zhengwentai Sun, Yanghong Zhou, Honghong He, P. Y. Mok

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by ACM MM'23

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15199 2023-08-16 cs.CV cs.AI cs.CL cs.LG 67%

PromptStyler: Prompt-driven Style Generation for Source-free Domain Generalization

Junhyeong Cho, Gilhyun Nam, Sungyeon Kim, Hunmin Yang, Suha Kwak

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ICCV 2023, Project Page: https://promptstyler.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01546 2023-08-04 cs.SD cs.AI cs.LG cs.MM eess.AS 67%

MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies

Ke Chen, Yusong Wu, Haohe Liu, Marianna Nezhurina, Taylor Berg-Kirkpatrick, Shlomo Dubnov

专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI、cs.MM、eess.AS

Comments 16 pages, 3 figures, 2 tables, demo page: https://musicldm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02971 2023-07-07 cs.CV cs.AI cs.CL 67%

On the Cultural Gap in Text-to-Image Generation

Bingshuai Liu, Longyue Wang, Chenyang Lyu, Yong Zhang, Jinsong Su, Shuming Shi, Zhaopeng Tu

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Equal contribution: Bingshuai Liu and Longyue Wang. Work done while Bingshuai Liu and Chengyang Lyu were interning at Tencent AI Lab. Zhaopeng Tu is the corresponding author

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07760 2023-05-29 cs.AI cs.CL cs.MM 67%

Unsupervised Melody-Guided Lyrics Generation

Yufei Tian, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Gunnar Sigurdsson, Chenyang Tao, Wenbo Zhao, Tagyoung Chung, Jing Huang, Nanyun Peng

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM

Comments Presented at AAAI23 CreativeAI workshop (Non-Archival). A later version is accepted to ACL23

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.04604 2023-04-26 cs.RO cs.AI cs.CL cs.CV cs.LG 67%

StructDiffusion: Language-Guided Creation of Physically-Valid Structures using Unseen Objects

Weiyu Liu, Yilun Du, Tucker Hermans, Sonia Chernova, Chris Paxton

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to Robotics: Science and Systems (RSS) 2023. The previous version appeared in CoRL Workshop on Language and Robot Learning 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14794 2022-12-12 cs.CV cs.AI cs.LG cs.MM stat.ML 67%

Traditional Classification Neural Networks are Good Generators: They are Competitive with DDPMs and GANs

Guangrun Wang, Philip H. S. Torr

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI、cs.MM

Comments This paper has 29 pages with 22 figures, including rich supplementary information. Project page is at \url{https://classifier-as-generator.github.io/}

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13729 2022-10-26 cs.AI cs.CL cs.CV 67%

Hybrid Reinforced Medical Report Generation with M-Linear Attention and Repetition Penalty

Wenting Xu, Zhenghua Xu, Junyang Chen, Chang Qi, Thomas Lukasiewicz

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments This paper is current under peer-review in IEEE TNNLS

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11253 2022-08-25 cs.CV cs.AI cs.CL cs.LG 67%

FashionVQA: A Domain-Specific Visual Question Answering System

Min Wang, Ata Mahjoubfar, Anupama Joshi

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.09052 2022-03-18 cs.CV cs.AI cs.CL 67%

DU-VLG: Unifying Vision-and-Language Generation via Dual Sequence-to-Sequence Pre-training

Luyang Huang, Guocheng Niu, Jiachen Liu, Xinyan Xiao, Hua Wu

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments To appear at Findings of ACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.09756 2021-10-20 cs.CV cs.CL cs.MM 67%

A Picture is Worth a Thousand Words: A Unified System for Diverse Captions and Rich Images Generation

Yupan Huang, Bei Liu, Jianlong Fu, Yutong Lu

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments ACM MM 2021 (Video and Demo Track). Code: https://github.com/researchmm/generate-it

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.11240 2021-09-01 cs.CL cs.CV cs.MM 67%

HUMBO: Bridging Response Generation and Facial Expression Synthesis

Shang-Yu Su, Po-Wei Lin, Yun-Nung Chen

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments The first two authors contributed to this work equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.10876 2021-06-22 cs.CV cs.AI cs.MM 67%

Total Generate: Cycle in Cycle Generative Adversarial Networks for Generating Human Faces, Hands, Bodies, and Natural Scenes

Hao Tang, Nicu Sebe

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted to TMM, an extended version of a paper published in ACM MM 2019. arXiv admin note: substantial text overlap with arXiv:1908.00999

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.02779 2021-05-25 cs.CL cs.AI cs.CV cs.LG 67%

Unifying Vision-and-Language Tasks via Text Generation

Jaemin Cho, Jie Lei, Hao Tan, Mohit Bansal

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICML 2021 (15 pages, 4 figures, 14 tables)

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.09676 2020-07-03 cs.CV cs.AI cs.CL 67%

CORAL8: Concurrent Object Regression for Area Localization in Medical Image Panels

Sam Maksoud, Arnold Wiliem, Kun Zhao, Teng Zhang, Lin Wu, Brian C. Lovell

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted for MICCAI 2019

详情

展开后加载摘要…

URL PDF HTML 收藏