arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4965 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4965 篇

2204.14217 2022-05-30 cs.CV cs.LG 57%

CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers

Ming Ding, Wendi Zheng, Wenyi Hong, Jie Tang

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.11487 2022-05-24 cs.CV cs.LG 57%

Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, Mohammad Norouzi

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.09376 2022-05-20 cs.AR cs.AI cs.LG 57%

Multi-DNN Accelerators for Next-Generation AI Systems

Stylianos I. Venieris, Christos-Savvas Bouganis, Nicholas D. Lane

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments Accepted for publication at the IEEE Computer journal, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.06522 2022-05-16 cs.CL 57%

Joint Generation of Captions and Subtitles with Dual Decoding

Jitao Xu, François Buet, Josep Crego, Elise Bertin-Lemée, François Yvon

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CL

Comments Accepted at IWSLT 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.00602 2022-05-11 cs.CV 57%

Pro-UIGAN: Progressive Face Hallucination from Occluded Thumbnails

Yang Zhang, Xin Yu, Xiaobo Lu, Ping Liu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.02540 2022-05-06 cs.GR cs.CV 57%

Real-time Controllable Motion Transition for Characters

Xiangjun Tang, He Wang, Bo Hu, Xu Gong, Ruifan Yi, Qilong Kou, Xiaogang Jin

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Journal ref ACM Transactions on Graphics (Proc. Siggraph 2022), 2022, 41(4)

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.03175 2022-04-29 cs.CL 57%

A Survey on Dialogue Summarization: Recent Advances and New Frontiers

Xiachong Feng, Xiaocheng Feng, Bing Qin

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

Comments IJCAI 2022 Survey Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.01955 2022-04-06 cs.CV 57%

Autoregressive 3D Shape Generation via Canonical Mapping

An-Chieh Cheng, Xueting Li, Sifei Liu, Min Sun, Ming-Hsuan Yang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00792 2022-04-05 cs.CV 57%

IR-GAN: Image Manipulation with Linguistic Instruction by Increment Reasoning

Zhenhuan Liu, Jincan Deng, Liang Li, Shaofei Cai, Qianqian Xu, Shuhui Wang, Qingming Huang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Journal ref Proceedings of the 28th ACM International Conference on Multimedia,2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.14681 2022-03-31 cs.CV 57%

ObjectFormer for Image Manipulation Detection and Localization

Junke Wang, Zuxuan Wu, Jingjing Chen, Xintong Han, Abhinav Shrivastava, Ser-Nam Lim, Yu-Gang Jiang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments accepted by CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15799 2022-03-30 cs.CV 57%

StyleT2I: Toward Compositional and High-Fidelity Text-to-Image Synthesis

Zhiheng Li, Martin Renqiang Min, Kai Li, Chenliang Xu

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15709 2022-03-30 cs.CV 57%

OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object Interaction

Lixin Yang, Kailin Li, Xinyu Zhan, Fei Wu, Anran Xu, Liu Liu, Cewu Lu

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15334 2022-03-30 cs.CV 57%

AnyFace: Free-style Text-to-Face Synthesis and Manipulation

Jianxin Sun, Qiyao Deng, Qi Li, Muyi Sun, Min Ren, Zhenan Sun

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.11796 2022-03-22 cs.CV 57%

DOC2PPT: Automatic Presentation Slides Generation from Scientific Documents

Tsu-Jui Fu, William Yang Wang, Daniel McDuff, Yale Song

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments AAAI'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.06605 2022-03-16 cs.CV 57%

Depth-Aware Generative Adversarial Network for Talking Head Video Generation

Fa-Ting Hong, Longhao Zhang, Li Shen, Dan Xu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments 15 Pages; Accepted by CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.06107 2022-03-14 cs.CV 57%

REX: Reasoning-aware and Grounded Explanation

Shi Chen, Qi Zhao

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments To appear in CVPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03963 2022-03-14 cs.CV 57%

InfinityGAN: Towards Infinite-Pixel Image Synthesis

Chieh Hubert Lin, Hsin-Ying Lee, Yen-Chi Cheng, Sergey Tulyakov, Ming-Hsuan Yang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ICLR 2022. Full Paper: https://openreview.net/forum?id=ufGMqIM0a4b ; Project page: https://hubert0527.github.io/infinityGAN/

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.12492 2022-03-07 cs.CV 57%

ISF-GAN: An Implicit Style Function for High-Resolution Image-to-Image Translation

Yahui Liu, Yajing Chen, Linchao Bao, Nicu Sebe, Bruno Lepri, Marco De Nadai

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments 14 pages, 15 figures

Journal ref IEEE Transactions on Multimedia, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.12211 2022-02-25 cs.CV 57%

Self-Distilled StyleGAN: Towards Generation from Internet Photos

Ron Mokady, Michal Yarom, Omer Tov, Oran Lang, Daniel Cohen-Or, Tali Dekel, Michal Irani, Inbar Mosseri

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.08124 2022-02-17 cs.CL 57%

XFBoost: Improving Text Generation with Controllable Decoders

Xiangyu Peng, Michael Sollami

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.10887 2022-02-04 cs.CV 57%

Trust It or Not: Confidence-Guided Automatic Radiology Report Generation

Yixin Wang, Zihao Lin, Zhe Xu, Haoyu Dong, Jiang Tian, Jie Luo, Zhongchao Shi, Yang Zhang, Jianping Fan, Zhiqiang He

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.01111 2022-01-10 cs.CV 57%

A Comprehensive Survey of Scene Graphs: Generation and Application

Xiaojun Chang, Pengzhen Ren, Pengfei Xu, Zhihui Li, Xiaojiang Chen, Alex Hauptmann

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments 25 pages

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.06883 2021-12-14 cs.SE cs.AI cs.DC 57%

A Methodology for a Scalable, Collaborative, and Resource-Efficient Platform to Facilitate Healthcare AI Research

Raphael Y. Cohen, Vesela P. Kovacheva

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03517 2021-12-08 cs.CV cs.GR 57%

CG-NeRF: Conditional Generative Neural Radiance Fields

Kyungmin Jo, Gyumin Shim, Sanghun Jung, Soyoung Yang, Jaegul Choo

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00180 2021-12-02 cs.CV cs.GR 57%

SpaceEdit: Learning a Unified Editing Space for Open-Domain Image Editing

Jing Shi, Ning Xu, Haitian Zheng, Alex Smith, Jiebo Luo, Chenliang Xu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.04556 2021-11-30 cs.CV 57%

Text as Neural Operator: Image Manipulation by Text Instruction

Tianhao Zhang, Hung-Yu Tseng, Lu Jiang, Weilong Yang, Honglak Lee, Irfan Essa

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.00484 2021-11-24 cs.CR cs.LG cs.SD eess.AS eess.IV 57%

Deepfakes Generation and Detection: State-of-the-art, open challenges, countermeasures, and way forward

Momina Masood, Marriam Nawaz, Khalid Mahmood Malik, Ali Javed, Aun Irtaza

专题命中 多模态生成 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.04735 2021-11-11 eess.IV cs.CV physics.med-ph 57%

Feature-enhanced Generation and Multi-modality Fusion based Deep Neural Network for Brain Tumor Segmentation with Missing MR Modalities

Tongxue Zhou, Stéphane Canu, Pierre Vera, Su Ruan

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 30 pages, 7 figures

Journal ref Neurocomputing 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.05527 2021-11-11 cs.AI 57%

LUMINOUS: Indoor Scene Generation for Embodied AI Challenges

Yizhou Zhao, Kaixiang Lin, Zhiwei Jia, Qiaozi Gao, Govind Thattai, Jesse Thomason, Gaurav S. Sukhatme

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 2021 paper, Amazon

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.13290 2021-11-08 cs.CV cs.LG 57%

CogView: Mastering Text-to-Image Generation via Transformers

Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, Jie Tang

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments to appear in NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏