arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2402.01858 2024-04-19 cs.LG cs.AI cs.CL cs.CV 82%

Explaining latent representations of generative models with large multimodal models

Mengdan Zhu, Zhenke Liu, Bo Pan, Abhinav Angirekula, Liang Zhao

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICLR 2024 Workshop on Reliable and Responsible Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14652 2024-03-25 cs.CY cs.AI cs.CL cs.MM 82%

MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation

Han Wang, Roy Ka-Wei Lee

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments 8 pages, 7 figures, ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03040 2024-02-06 cs.CV cs.AI cs.LG cs.MM 82%

InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions

Yiyuan Zhang, Yuhao Kang, Zhixin Zhang, Xiaohan Ding, Sanyuan Zhao, Xiangyu Yue

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Code, models, and demo are available at https://github.com/invictus717/InteractiveVideo

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02369 2024-02-06 cs.CV cs.CL cs.MM 82%

M$^3$Face: A Unified Multi-Modal Multilingual Framework for Human Face Generation and Editing

Mohammadreza Mofayezi, Reza Alipour, Mohammad Ali Kakavand, Ehsaneddin Asgari

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01504 2023-12-05 cs.CV cs.AI cs.CL cs.LG 82%

Effectively Fine-tune to Improve Large Multimodal Models for Radiology Report Generation

Yuzhe Lu, Sungmin Hong, Yash Shah, Panpan Xu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to Deep Generative Models for Health Workshop at NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11090 2023-11-21 cs.CV cs.AI cs.CL cs.LG 82%

Beyond Images: An Integrative Multi-modal Approach to Chest X-Ray Report Generation

Nurbanu Aksoy, Serge Sharoff, Selcuk Baser, Nishant Ravikumar, Alejandro F Frangi

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.17842 2023-10-31 cs.CV cs.CL cs.MM 82%

SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs

Lijun Yu, Yong Cheng, Zhiruo Wang, Vivek Kumar, Wolfgang Macherey, Yanping Huang, David A. Ross, Irfan Essa, Yonatan Bisk, Ming-Hsuan Yang, Kevin Murphy, Alexander G. Hauptmann, Lu Jiang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments NeurIPS 2023 spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10765 2023-10-24 cs.CV cs.AI cs.CL 82%

BiomedJourney: Counterfactual Biomedical Image Generation by Instruction-Learning from Multimodal Patient Journeys

Yu Gu, Jianwei Yang, Naoto Usuyama, Chunyuan Li, Sheng Zhang, Matthew P. Lungren, Jianfeng Gao, Hoifung Poon

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project page & demo: https://aka.ms/biomedjourney

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13361 2023-10-23 cs.CV cs.AI cs.CL 82%

Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine Translation

Wenyu Guo, Qingkai Fang, Dong Yu, Yang Feng

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2023 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02051 2023-08-24 cs.CV cs.AI cs.MM 82%

Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing

Alberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia, Marco Bertini, Rita Cucchiara

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10843 2023-08-22 cs.MM cs.CV cs.LG cs.SD eess.AS 82%

TranSTYLer: Multimodal Behavioral Style Transfer for Facial and Body Gestures Generation

Mireille Fares, Catherine Pelachaud, Nicolas Obin

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11846 2023-05-22 cs.CV cs.CL cs.LG cs.SD eess.AS 82%

Any-to-Any Generation via Composable Diffusion

Zineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng, Mohit Bansal

专题命中 多模态生成 :any-to-any(title);multimodal(abstract);分类 cs.CV、cs.CL、eess.AS

Comments Project Page: https://codi-gen.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11832 2023-05-22 stat.ML cs.LG 82%

Improving Multimodal Joint Variational Autoencoders through Normalizing Flows and Correlation Analysis

Agathe Senellart, Clément Chadebec, Stéphanie Allassonnière

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06756 2023-03-31 cs.CV cs.AI cs.MM cs.NE 82%

Decoding Visual Neural Representations by Multimodal Learning of Brain-Visual-Linguistic Features

Changde Du, Kaicheng Fu, Jinpeng Li, Huiguang He

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08672 2023-02-20 cs.LG cs.AI cs.CL cs.CV 82%

Multimodal Subtask Graph Generation from Instructional Videos

Yunseok Jang, Sungryull Sohn, Lajanugen Logeswaran, Tiange Luo, Moontae Lee, Honglak Lee

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.13235 2022-11-28 cs.CL cs.AI cs.CV 82%

Unified Multimodal Model with Unlikelihood Training for Visual Dialog

Zihao Wang, Junli Wang, Changjun Jiang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by the 30th ACM International Conference on Multimedia (ACM MM 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02127 2022-07-06 cs.LG stat.ML 82%

A survey of multimodal deep generative models

Masahiro Suzuki, Yutaka Matsuo

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

Comments Published in Advanced Robotics

Journal ref Advanced Robotics, 36:5-6, 261-278, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.11705 2022-05-25 cs.CV cs.AI cs.MM 82%

M6-Fashion: High-Fidelity Multi-modal Image Generation and Editing

Zhikang Li, Huiling Zhou, Shuai Bai, Peike Li, Chang Zhou, Hongxia Yang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments arXiv admin note: text overlap with arXiv:2105.14211

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03534 2022-05-10 cs.CL cs.CV cs.MM 82%

Attract me to Buy: Advertisement Copywriting Generation with Multimodal Multi-structured Information

Zhipeng Zhang, Xinglin Hou, Kai Niu, Zhongzhen Huang, Tiezheng Ge, Yuning Jiang, Qi Wu, Peng Wang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.00590 2022-03-29 cs.CL cs.AI cs.CV cs.LG 82%

WebQA: Multihop and Multimodal QA

Yingshan Chang, Mridu Narang, Hisami Suzuki, Guihong Cao, Jianfeng Gao, Yonatan Bisk

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments CVPR Camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.09753 2021-10-20 cs.CV cs.CL cs.MM 82%

Unifying Multimodal Transformer for Bi-directional Image and Text Generation

Yupan Huang, Hongwei Xue, Bei Liu, Yutong Lu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments ACM MM 2021 (Industrial Track). Code: https://github.com/researchmm/generate-it

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.08478 2021-09-20 cs.CL cs.CV cs.MM 82%

Multimodal Incremental Transformer with Visual Grounding for Visual Dialogue Generation

Feilong Chen, Fandong Meng, Xiuyi Chen, Peng Li, Jie Zhou

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments ACL Fingdings 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.04324 2021-09-17 cs.CL cs.AI cs.CV 82%

FairyTailor: A Multimodal Generative Framework for Storytelling

Eden Bensaid, Mauro Martino, Benjamin Hoover, Hendrik Strobelt

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments visit https://fairytailor.org/ and https://github.com/EdenBD/MultiModalStory-demo for web demo and source code

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.09252 2020-05-20 cs.IR 82%

Multi-Modal Summary Generation using Multi-Objective Optimization

Anubhav Jangra, Sriparna Saha, Adam Jatowt, Mohammad Hasanuzzaman

专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract)

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.01707 2020-01-07 cs.LG eess.IV stat.ML 82%

Meta-modal Information Flow: A Method for Capturing Multimodal Modular Disconnectivity in Schizophrenia

Haleh Falakshahi, Victor M. Vergara, Jingyu Liu, Daniel H. Mathalon, Judith M. Ford, James Voyvodic, Bryon A. Mueller, Aysenil Belger, Sarah McEwen, Steven G. Potkin, Adrian Preda, Hooman Rokham, Jing Sui, Jessica A. Turner, Sergey Plis, Vince D. Calhoun

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

Journal ref IEEE Transactions on Biomedical Engineering, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.03393 2019-11-11 stat.ML cs.LG 82%

Variational Mixture-of-Experts Autoencoders for Multi-Modal Deep Generative Models

Yuge Shi, N. Siddharth, Brooks Paige, Philip H. S. Torr

专题命中 多模态生成 :multi-modal(title,abstract);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.03986 2019-10-18 cs.CL cs.AI cs.CV 82%

Multimodal Differential Network for Visual Question Generation

Badri N. Patro, Sandeep Kumar, Vinod K. Kurmi, Vinay P. Namboodiri

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments EMNLP 2018 (accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.08129 2018-02-23 cs.AI cs.CL cs.CV 82%

Multimodal Explanations: Justifying Decisions and Pointing to the Evidence

Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, Marcus Rohrbach

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments arXiv admin note: text overlap with arXiv:1612.04757

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17585 2026-04-21 cs.CV cs.AI cs.LG 82%

DGSSM: Diffusion guided state-space models for multimodal salient object detection

DGSSM:基于扩散引导的状态空间模型的多模态显著目标检测

Suklav Ghosh, Arijit Sur, Pinaki Mitra

机构 * Dept. of Computer Science and Engineering, Indian Institute of Technology, Guwahati(计算机科学与工程系,印度理工学院,果阿提)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出DGSSM,一种结合扩散模型结构先验和多尺度状态空间编码的多模态显著目标检测框架,通过迭代Mamba扩散细化机制提升边界精度,实验表明其在多个评估指标上优于现有方法。

Comments Accepted at ICPR 2026. Diffusion-guided Mamba framework for multimodal salient object detection. Evaluated on 13 benchmarks (RGB, RGB-D, RGB-T)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15512 2024-10-23 cs.CV cs.AI 82%

PixelBytes: Catching Unified Embedding for Multimodal Generation

Fabien Furfaro

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments This article is an earlier version of my work arXiv:2410.01820 "PixelBytes: Catching Unified Representation for Multimodal Generation."

详情

展开后加载摘要…

URL PDF HTML 收藏