arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4979 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4979 篇

2303.02688 2023-03-09 cs.CV 79%

Text2Face: A Multi-Modal 3D Face Model

Will Rowan, Patrik Huber, Nick Pears, Andrew Keeling

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Fixed formatting and a typo

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.00751 2023-01-18 cs.CV 79%

CSDN: Cross-modal Shape-transfer Dual-refinement Network for Point Cloud Completion

Zhe Zhu, Liangliang Nan, Haoran Xie, Honghua Chen, Mingqiang Wei, Jun Wang, Jing Qin

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.02936 2022-12-08 cs.CV 79%

M-VADER: A Model for Diffusion with Multimodal Context

Samuel Weinbach, Marco Bellagente, Constantin Eichenberg, Andrew Dai, Robert Baldock, Souradeep Nanda, Björn Deiseroth, Koen Oostermeijer, Hannah Teufel, Andres Felipe Cruz-Salinas

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments 22 pages, 14 figures, 2 tables, fixed figure 3

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12565 2022-10-25 cs.CL cs.LG 79%

A Visual Tour Of Current Challenges In Multimodal Language Models

Shashank Sonkar, Naiming Liu, Richard G. Baraniuk

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13360 2022-09-29 cs.CV 79%

Draw Your Art Dream: Diverse Digital Art Synthesis with Multimodal Guided Diffusion

Nisha Huang, Fan Tang, Weiming Dong, Changsheng Xu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.05490 2022-09-29 cs.AI 79%

Modeling and solving the multimodal car- and ride-sharing problem

Miriam Enzi, Sophie N. Parragh, David Pisinger, Matthias Prandtstetter

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.12998 2022-09-28 cs.LG cs.AI cs.CY 79%

Integrated multimodal artificial intelligence framework for healthcare applications

Luis R. Soenksen, Yu Ma, Cynthia Zeng, Leonard D. J. Boussioux, Kimberly Villalobos Carballo, Liangyuan Na, Holly M. Wiberg, Michael L. Li, Ignacio Fuentes, Dimitris Bertsimas

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Journal ref Nature npj Digital Medicine, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.03160 2022-09-09 cs.CV 79%

AI Illustrator: Translating Raw Descriptions into Images by Prompt-based Cross-Modal Generation

Yiyang Ma, Huan Yang, Bei Liu, Jianlong Fu, Jiaying Liu

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.02399 2022-09-07 cs.LG cs.CV eess.IV 79%

Multi-Modal Hypergraph Diffusion Network with Dual Prior for Alzheimer Classification

Angelica I. Aviles-Rivero, Christina Runkel, Nicolas Papadakis, Zoe Kourtzi, Carola-Bibiane Schönlieb

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Journal ref MICCAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.04854 2022-08-03 cs.CV 79%

Identity-guided Face Generation with Multi-modal Contour Conditions

Qingyan Bai, Weihao Xia, Fei Yin, Yujiu Yang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to ICIP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.11795 2022-07-26 cs.CV 79%

Cross-Modal 3D Shape Generation and Manipulation

Zezhou Cheng, Menglei Chai, Jian Ren, Hsin-Ying Lee, Kyle Olszewski, Zeng Huang, Subhransu Maji, Sergey Tulyakov

专题命中 多模态生成 :cross-modal(title);multi-modal(abstract);分类 cs.CV

Comments ECCV 2022. Project page: https://people.cs.umass.edu/~zezhoucheng/edit3d/

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.04713 2022-07-12 cs.CL 79%

GMN: Generative Multi-modal Network for Practical Document Information Extraction

Haoyu Cao, Jiefeng Ma, Antai Guo, Yiqing Hu, Hao Liu, Deqiang Jiang, Yinsong Liu, Bo Ren

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted to NAACL 2022 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.08970 2022-06-22 cs.CV 79%

MultiEarth 2022 -- The Champion Solution for the Matrix Completion Challenge via Multimodal Regression and Generation

Bo Peng, Hongchen Liu, Hang Zhou, Yuchuan Gou, Jui-Hsin Lai

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2022, MultiEarth 2022, Matrix Completion Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.09808 2022-06-15 cs.CV 79%

Integrated Construction of Multimodal Atlases with Structural Connectomes in the Space of Riemannian Metrics

Kristen M. Campbell, Haocheng Dai, Zhe Su, Martin Bauer, P. Thomas Fletcher, Sarang C. Joshi

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) https://www.melba-journal.org/papers/2022:016.html. arXiv admin note: substantial text overlap with arXiv:2103.05730

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.13871 2022-06-14 cs.SD cs.GR cs.LG eess.AS 79%

Transflower: probabilistic autoregressive dance generation with multimodal attention

Guillermo Valle-Pérez, Gustav Eje Henter, Jonas Beskow, André Holzapfel, Pierre-Yves Oudeyer, Simon Alexanderson

专题命中 多模态生成 :multimodal(title,abstract);分类 eess.AS

Comments Article presented at SIGGRAPH Asia 2021, and published in ACM Transactions on Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.05039 2022-06-13 cs.CV 79%

Image Generation with Multimodal Priors using Denoising Diffusion Probabilistic Models

Nithin Gopalakrishnan Nair, Wele Gedara Chaminda Bandara, Vishal M Patel

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.01988 2022-06-07 cs.CV 79%

Cross-modal Clinical Graph Transformer for Ophthalmic Report Generation

Mingjie Li, Wenjia Cai, Karin Verspoor, Shirui Pan, Xiaodan Liang, Xiaojun Chang

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

Comments CVPR 2022 (Poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.13258 2022-04-29 cs.CL 79%

Cross-modal Memory Networks for Radiology Report Generation

Zhihong Chen, Yaling Shen, Yan Song, Xiang Wan

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CL

Comments Natural Language Processing. 11 pages, 6 figures. ACL-IJCNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12667 2022-04-28 cs.CV 79%

MM-TTA: Multi-Modal Test-Time Adaptation for 3D Semantic Segmentation

Inkyu Shin, Yi-Hsuan Tsai, Bingbing Zhuang, Samuel Schulter, Buyu Liu, Sparsh Garg, In So Kweon, Kuk-Jin Yoon

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.14211 2022-02-22 cs.CV 79%

M6-UFC: Unifying Multi-Modal Controls for Conditional Image Synthesis via Non-Autoregressive Generative Transformers

Zhu Zhang, Jianxin Ma, Chang Zhou, Rui Men, Zhikang Li, Ming Ding, Jie Tang, Jingren Zhou, Hongxia Yang

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS21

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03608 2022-02-01 cs.LG cs.AI 79%

How to Sense the World: Leveraging Hierarchy in Multimodal Perception for Robust Reinforcement Learning Agents

Miguel Vasco, Hang Yin, Francisco S. Melo, Ana Paiva

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at the International Conference on Autonomous Agents and MultiAgent Systems (AAMAS) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.09702 2021-10-22 cs.CL 79%

A non-hierarchical attention network with modality dropout for textual response generation in multimodal dialogue systems

Rongyi Sun, Borun Chen, Qingyu Zhou, Yinghui Li, YunBo Cao, Hai-Tao Zheng

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Submitted to ICASSP2022 (currently under review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.02401 2021-10-12 cs.CL 79%

Vision Guided Generative Pre-trained Language Models for Multimodal Abstractive Summarization

Tiezheng Yu, Wenliang Dai, Zihan Liu, Pascale Fung

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments Long Paper Accepted in EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01229 2021-09-06 cs.CL cs.LG 79%

Multimodal Conditionality for Natural Language Generation

Michael Sollami, Aashish Jain

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.09953 2021-07-22 cs.CV 79%

Characterization Multimodal Connectivity of Brain Network by Hypergraph GAN for Alzheimer's Disease Analysis

Junren Pan, Baiying Lei, Yanyan Shen, Yong Liu, Zhiguang Feng, Shuqiang Wang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.08673 2021-07-20 eess.IV cs.CV cs.LG 79%

Input Agnostic Deep Learning for Alzheimer's Disease Classification Using Multimodal MRI Images

Aidana Massalimova, Huseyin Atakan Varol

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments 4 pages, submitted to EMBC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.00419 2021-07-19 cs.CL 79%

KM-BART: Knowledge Enhanced Multimodal BART for Visual Commonsense Generation

Yiran Xing, Zai Shi, Zhao Meng, Gerhard Lakemeyer, Yunpu Ma, Roger Wattenhofer

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

Comments ACL-IJCNLP 2021 main conference. The first three authors contribute equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.15309 2021-06-30 cs.CV 79%

Multimodal Semantic Scene Graphs for Holistic Modeling of Surgical Procedures

Ege Özsoy, Evin Pınar Örnek, Ulrich Eck, Federico Tombari, Nassir Navab

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.14445 2021-06-01 cs.CL 79%

Modeling Text-visual Mutual Dependency for Multi-modal Dialog Generation

Shuhe Wang, Yuxian Meng, Xiaofei Sun, Fei Wu, Rongbin Ouyang, Rui Yan, Tianwei Zhang, Jiwei Li

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL

Comments arXiv admin note: text overlap with arXiv:2012.15015

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.07854 2021-03-23 cs.CV 79%

Three Steps to Multimodal Trajectory Prediction: Modality Clustering, Classification and Synthesis

Jianhua Sun, Yuxuan Li, Hao-Shu Fang, Cewu Lu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏