arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4965 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4965 篇

2406.09406 2024-06-17 cs.CV cs.AI cs.LG 81%

4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities

Roman Bachmann, Oğuzhan Fatih Kar, David Mizrahi, Ali Garjani, Mingfei Gao, David Griffiths, Jiaming Hu, Afshin Dehghan, Amir Zamir

专题命中 多模态生成 :any-to-any(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments Project page at 4m.epfl.ch

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06730 2024-06-12 cs.CV cs.AI 81%

TRINS: Towards Multimodal Language Models that Can Read

Ruiyi Zhang, Yanzhe Zhang, Jian Chen, Yufan Zhou, Jiuxiang Gu, Changyou Chen, Tong Sun

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20213 2024-05-31 cs.AI cs.CL cs.LG 81%

PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization

Vijay Jaisankar, Sambaran Bandyopadhyay, Kalp Vyas, Varre Chaitanya, Shwetha Somasundaram

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14545 2024-05-30 cs.CL cs.CV 81%

Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective

Zihao Yue, Liang Zhang, Qin Jin

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13195 2024-05-24 cs.CV cs.AI 81%

CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers

Andrew Marmon, Grant Schindler, José Lezama, Dan Kondratyuk, Bryan Seybold, Irfan Essa

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07553 2024-05-14 cs.AI cs.CL 81%

Hijacking Context in Large Multi-modal Models

Joonhyun Jeong

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments Technical Report. Preprint

Journal ref ICLR 2024 Workshop on Reliable and Responsible Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18591 2024-04-30 cs.CV cs.AI 81%

FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion

Abhishek Kumar Singh, Ioannis Patras

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15100 2024-04-24 cs.CV cs.MM 81%

Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation

Xun Wu, Shaohan Huang, Furu Wei

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13791 2024-04-23 cs.CV cs.AI 81%

Universal Fingerprint Generation: Controllable Diffusion Model with Multimodal Conditions

Steven A. Grosz, Anil K. Jain

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16749 2024-04-18 cs.CV cs.AI eess.IV 81%

MISC: Ultra-low Bitrate Image Semantic Compression Driven by Large Multimodal Model

Chunyi Li, Guo Lu, Donghui Feng, Haoning Wu, Zicheng Zhang, Xiaohong Liu, Guangtao Zhai, Weisi Lin, Wenjun Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 13 page, 11 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17546 2024-04-10 cs.CV cs.AI cs.LG 81%

PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor

Vidit Goel, Elia Peruzzo, Yifan Jiang, Dejia Xu, Xingqian Xu, Nicu Sebe, Trevor Darrell, Zhangyang Wang, Humphrey Shi

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in CVPR 2024, Project page https://vidit98.github.io/publication/conference-paper/pair_diff.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07362 2024-04-03 cs.CL cs.CV 81%

Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision

Seongyun Lee, Sue Hyun Park, Yongrae Jo, Minjoon Seo

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00588 2024-04-02 cs.CV cs.AI 81%

Memory-based Cross-modal Semantic Alignment Network for Radiology Report Generation

Yitian Tao, Liyan Ma, Jing Yu, Han Zhang

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15059 2024-03-25 cs.CV cs.AI 81%

MM-Diff: High-Fidelity Image Personalization via Multi-Modal Condition Integration

Zhichao Wei, Qingkun Su, Long Qin, Weizhi Wang

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04302 2024-03-22 cs.CV cs.CL 81%

Prompt Highlighter: Interactive Control for Multi-Modal LLMs

Yuechen Zhang, Shengju Qian, Bohao Peng, Shu Liu, Jiaya Jia

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments CVPR 2024; Project Page: https://julianjuaner.github.io/projects/PromptHighlighter

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04789 2024-03-12 cs.CL cs.AI cs.LG 81%

TopicDiff: A Topic-enriched Diffusion Approach for Multimodal Conversational Emotion Detection

Jiamin Luo, Jingjing Wang, Guodong Zhou

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13587 2024-03-08 cs.CL cs.CV 81%

A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation

Yunxin Li, Baotian Hu, Wenhan Luo, Lin Ma, Yuxin Ding, Min Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16117 2024-02-27 cs.RO cs.AI cs.CV 81%

RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis

Yao Mu, Junting Chen, Qinglong Zhang, Shoufa Chen, Qiaojun Yu, Chongjian Ge, Runjian Chen, Zhixuan Liang, Mengkang Hu, Chaofan Tao, Peize Sun, Haibao Yu, Chao Yang, Wenqi Shao, Wenhai Wang, Jifeng Dai, Yu Qiao, Mingyu Ding, Ping Luo

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07276 2024-01-30 cs.CL cs.AI cs.LG q-bio.BM 81%

BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations

Qizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu, Kaiyuan Gao, Lijun Wu, Yingce Xia, Rui Yan

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by Empirical Methods in Natural Language Processing 2023 (EMNLP 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13298 2024-01-25 cs.CL cs.AI 81%

Towards Explainable Harmful Meme Detection through Multimodal Debate between Large Language Models

Hongzhan Lin, Ziyang Luo, Wei Gao, Jing Ma, Bo Wang, Ruichao Yang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments The first work towards explainable harmful meme detection by harnessing advanced LLMs

Journal ref The ACM Web Conference 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11631 2024-01-23 cs.CV cs.CL cs.LG 81%

Text-to-Image Cross-Modal Generation: A Systematic Review

Maciej Żelaszczyk, Jacek Mańdziuk

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.05134 2024-01-11 cs.AI cs.CL 81%

Yes, this is what I was looking for! Towards Multi-modal Medical Consultation Concern Summary Generation

Abhisek Tiwari, Shreyangshu Bera, Sriparna Saha, Pushpak Bhattacharyya, Samrat Ghosh

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02433 2024-01-08 cs.CV cs.AI cs.LG 81%

FedDiff: Diffusion Model Driven Federated Learning for Multi-Modal and Multi-Clients

DaiXun Li, Weiying Xie, ZiXuan Wang, YiBing Lu, Yunsong Li, Leyuan Fang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15296 2023-12-21 cs.CV cs.AI cs.LG 81%

MultiFusion: Fusing Pre-Trained Models for Multi-Lingual, Multi-Modal Image Generation

Marco Bellagente, Manuel Brack, Hannah Teufel, Felix Friedrich, Björn Deiseroth, Constantin Eichenberg, Andrew Dai, Robert Baldock, Souradeep Nanda, Koen Oostermeijer, Andres Felipe Cruz-Salinas, Patrick Schramowski, Kristian Kersting, Samuel Weinbach

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments Proceedings of Advances in Neural Information Processing Systems: Annual Conference on Neural Information Processing Systems (NeurIPS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09767 2023-12-20 cs.CL cs.AI 81%

VLIS: Unimodal Language Models Guide Multimodal Language Generation

Jiwan Chung, Youngjae Yu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted as main paper in EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06647 2023-12-12 cs.CV cs.AI cs.LG 81%

4M: Massively Multimodal Masked Modeling

David Mizrahi, Roman Bachmann, Oğuzhan Fatih Kar, Teresa Yeo, Mingfei Gao, Afshin Dehghan, Amir Zamir

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2023 Spotlight. Project page at https://4m.epfl.ch/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09909 2023-12-05 cs.CV cs.CL 81%

Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis

Chaoyi Wu, Jiayu Lei, Qiaoyu Zheng, Weike Zhao, Weixiong Lin, Xiaoman Zhang, Xiao Zhou, Ziheng Zhao, Ya Zhang, Yanfeng Wang, Weidi Xie

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16483 2023-11-29 cs.CV cs.CL 81%

ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Yucheng Han, Chi Zhang, Xin Chen, Xu Yang, Zhibin Wang, Gang Yu, Bin Fu, Hanwang Zhang

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL

Comments Code and model on https://tingxueronghua.github.io/ChartLlama/

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10056 2023-11-03 cs.CV cs.MM 81%

GlueGen: Plug and Play Multi-modal Encoders for X-to-image Generation

Can Qin, Ning Yu, Chen Xing, Shu Zhang, Zeyuan Chen, Stefano Ermon, Yun Fu, Caiming Xiong, Ran Xu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15585 2023-10-25 cs.CL cs.CV cs.LG 81%

Multimodal Representations for Teacher-Guided Compositional Visual Reasoning

Wafa Aissa, Marin Ferecatu, Michel Crucianu

专题命中 多模态生成 :multimodal(title);cross-modal(abstract);分类 cs.CV、cs.CL

Journal ref Advanced Concepts for Intelligent Vision Systems, 21st International Conference (ACIVS 2023), Aug 2023, Kumamoto, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏