arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4965 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4965 篇

2301.10799 2023-02-14 cs.CL 57%

Towards a Unified Model for Generating Answers and Explanations in Visual Question Answering

Chenxi Whitehouse, Tillman Weyde, Pranava Madhyastha

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments Findings of EACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.15480 2023-02-03 cs.LG cs.AI stat.ML 57%

Post-hoc Concept Bottleneck Models

Mert Yuksekgonul, Maggie Wang, James Zou

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments ICLR 2023 Spotlight (notable-top-25%)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.10172 2023-01-31 cs.CL cs.LG 57%

MTTN: Multi-Pair Text to Text Narratives for Prompt Generation

Archan Ghosh, Debgandhar Ghosh, Madhurima Maji, Suchinta Chanda, Kalporup Goswami

专题命中 多模态生成 :image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.05402 2023-01-16 cs.CL cs.LG 57%

In BLOOM: Creativity and Affinity in Artificial Lyrics and Art

Evan Crothers, Herna Viktor, Nathalie Japkowicz

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments Accepted to AAAI2023 creativeAI workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15977 2023-01-04 cs.CV 57%

One is All: Bridging the Gap Between Neural Radiance Fields Architectures with Progressive Volume Distillation

Shuangkang Fang, Weixin Xu, Heng Wang, Yi Yang, Yufeng Wang, Shuchang Zhou

专题命中 多模态生成 :any-to-any(abstract);分类 cs.CV

Comments Accepted by AAAI2023. Project Page: https://sk-fun.fun/PVD

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.09329 2022-12-20 cs.CV 57%

SrTR: Self-reasoning Transformer with Visual-linguistic Knowledge for Scene Graph Generation

Yuxiang Zhang, Zhenbo Liu, Shuai Wang

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.09248 2022-12-20 cs.CL cs.SE 57%

Natural Language to Code Generation in Interactive Data Science Notebooks

Pengcheng Yin, Wen-Ding Li, Kefan Xiao, Abhishek Rao, Yeming Wen, Kensen Shi, Joshua Howland, Paige Bailey, Michele Catasta, Henryk Michalewski, Alex Polozov, Charles Sutton

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

Comments 46 pages. 32 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.06908 2022-12-15 cs.NI cs.AI cs.IT cs.LG math.IT 57%

Enabling the Wireless Metaverse via Semantic Multiverse Communication

Jihong Park, Jinho Choi, Seong-Lyun Kim, Mehdi Bennis

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments 7 pages, 6 figures, submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.06365 2022-12-14 eess.SP eess.AS eess.IV 57%

Towards deep generation of guided wave representations for composite materials

Mahindra Rautela, J. Senthilnath, Armin Huber, S. Gopalakrishnan

专题命中 多模态生成 :multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.02746 2022-12-07 cs.AI cs.LG 57%

UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

Jiaqi Chen, Tong Li, Jinghui Qin, Pan Lu, Liang Lin, Chongyu Chen, Xiaodan Liang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.05744 2022-12-06 cs.CV cs.GR 57%

More Control for Free! Image Synthesis with Semantic Diffusion Guidance

Xihui Liu, Dong Huk Park, Samaneh Azadi, Gong Zhang, Arman Chopikyan, Yuxiao Hu, Humphrey Shi, Anna Rohrbach, Trevor Darrell

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments WACV 2023. Project page https://xh-liu.github.io/sdg/

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14842 2022-11-29 cs.CV 57%

Unified Discrete Diffusion for Simultaneous Vision-Language Generation

Minghui Hu, Chuanxia Zheng, Heliang Zheng, Tat-Jen Cham, Chaoyue Wang, Zuopeng Yang, Dacheng Tao, Ponnuthurai N. Suganthan

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07363 2022-11-15 q-bio.NC cs.AI cs.LG 57%

Joint Graph Convolution for Analyzing Brain Structural and Functional Connectome

Yueting Li, Qingyue Wei, Ehsan Adeli, Kilian M. Pohl, Qingyu Zhao

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.05865 2022-10-18 cs.CV 57%

DF-GAN: A Simple and Effective Baseline for Text-to-Image Synthesis

Ming Tao, Hao Tang, Fei Wu, Xiao-Yuan Jing, Bing-Kun Bao, Changsheng Xu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06978 2022-10-14 cs.CV cs.LG stat.ML 57%

LION: Latent Point Diffusion Models for 3D Shape Generation

Xiaohui Zeng, Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, Karsten Kreis

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.01813 2022-10-11 cs.CV 57%

TAG: Boosting Text-VQA via Text-aware Visual Question-answer Generation

Jun Wang, Mingfei Gao, Yuqian Hu, Ramprasaath R. Selvaraju, Chetan Ramaiah, Ran Xu, Joseph F. JaJa, Larry S. Davis

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments BMVC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.14988 2022-09-30 cs.CV cs.LG stat.ML 57%

DreamFusion: Text-to-3D using 2D Diffusion

Ben Poole, Ajay Jain, Jonathan T. Barron, Ben Mildenhall

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments see project page at https://dreamfusion3d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.14046 2022-09-29 cs.CV 57%

Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image Generation

Xintian Wu, Hanbin Zhao, Liangli Zheng, Shouhong Ding, Xi Li

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08583 2022-09-07 cs.CV 57%

VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance

Katherine Crowson, Stella Biderman, Daniel Kornis, Dashiell Stander, Eric Hallahan, Louis Castricato, Edward Raff

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted for publication at ECCV 2022 Code available at https://github.com/EleutherAI/vqgan-clip/tree/main/notebooks

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12675 2022-09-02 cs.CV 57%

Adaptively-Realistic Image Generation from Stroke and Sketch with Diffusion Model

Shin-I Cheng, Yu-Jie Chen, Wei-Chen Chiu, Hung-Yu Tseng, Hsin-Ying Lee

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11844 2022-08-26 cs.CV 57%

Unbiased Multi-Modality Guidance for Image Inpainting

Yongsheng Yu, Dawei Du, Libo Zhang, Tiejian Luo

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11546 2022-08-25 cs.CV 57%

Unsupervised Structure-Consistent Image-to-Image Translation

Shima Shahfar, Charalambos Poullis

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments structure-consistent image-to-image translation \and style transfer \and training class imbalance

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.08493 2022-08-19 cs.CV 57%

Text-to-Image Generation via Implicit Visual Guidance and Hypernetwork

Xin Yuan, Zhe Lin, Jason Kuen, Jianming Zhang, John Collomosse

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.07943 2022-08-18 cs.CV 57%

TRoVE: Transforming Road Scene Datasets into Photorealistic Virtual Environments

Shubham Dokania, Anbumani Subramanian, Manmohan Chandraker, C. V. Jawahar

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments 18 pages, 5 figures, Accepted in European Conference on Computer Vision (ECCV 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.04908 2022-08-09 cs.CV 57%

No Token Left Behind: Explainability-Aided Image Classification and Generation

Roni Paiss, Hila Chefer, Lior Wolf

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.11046 2022-08-02 cs.CV 57%

FRT-PAD: Effective Presentation Attack Detection Driven by Face Related Task

Wentian Zhang, Haozhe Liu, Feng Liu, Raghavendra Ramachandra, Christoph Busch

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.05776 2022-07-25 cs.CV 57%

Learning to Predict Diverse Human Motions from a Single Image via Mixture Density Networks

Chunzhi Gu, Yan Zhao, Chao Zhang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07165 2022-07-18 cs.SI cs.MM 57%

Estimating Emotion Contagion on Social Media via Localized Diffusion in Dynamic Graphs

Trisha Mittal, Puneet Mathur, Rohan Chandra, Apurva Bhatt, Vikram Gupta, Debdoot Mukherjee, Aniket Bera, Dinesh Manocha

专题命中 多模态生成 :audio-visual(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.06604 2022-07-15 cs.CV 57%

Rethinking Super-Resolution as Text-Guided Details Generation

Chenxi Ma, Bo Yan, Qing Lin, Weimin Tan, Siming Chen

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 11 figures, ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09114 2022-06-23 cs.CV cs.LG 57%

Bear the Query in Mind: Visual Grounding with Query-conditioned Convolution

Chonghan Chen, Qi Jiang, Chih-Hao Wang, Noel Chen, Haohan Wang, Xiang Li, Bhiksha Raj

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏