arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4965 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4965 篇

2310.00362 2023-10-03 cs.CV cs.GR 57%

Diffusion Posterior Illumination for Ambiguity-aware Inverse Rendering

Linjie Lyu, Ayush Tewari, Marc Habermann, Shunsuke Saito, Michael Zollhöfer, Thomas Leimkühler, Christian Theobalt

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments SIGGRAPH Asia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15807 2023-09-28 cs.CV 57%

Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Xiaoliang Dai, Ji Hou, Chih-Yao Ma, Sam Tsai, Jialiang Wang, Rui Wang, Peizhao Zhang, Simon Vandenhende, Xiaofang Wang, Abhimanyu Dubey, Matthew Yu, Abhishek Kadian, Filip Radenovic, Dhruv Mahajan, Kunpeng Li, Yue Zhao, Vladan Petrovic, Mitesh Kumar Singh, Simran Motwani, Yi Wen, Yiwen Song, Roshan Sumbaly, Vignesh Ramanathan, Zijian He, Peter Vajda, Devi Parikh

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15117 2023-09-27 cs.CV 57%

Generating Visual Scenes from Touch

Fengyu Yang, Jiacheng Zhang, Andrew Owens

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments ICCV 2023; Project site: https://fredfyyang.github.io/vision-from-touch/

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15988 2023-09-25 cs.CV cs.LG 57%

RGB-D-Fusion: Image Conditioned Depth Diffusion of Humanoid Subjects

Sascha Kirch, Valeria Olyunina, Jan Ondřej, Rafael Pagés, Sergio Martin, Clara Pérez-Molina

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11923 2023-09-22 cs.CV 57%

TextCLIP: Text-Guided Face Image Generation And Manipulation Without Adversarial Training

Xiaozhou You, Jian Zhang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07929 2023-09-22 cs.CV cs.LG 57%

Fast Adaptation with Bradley-Terry Preference Models in Text-To-Image Classification and Generation

Victor Gallego

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted to Proceedings of the 23rd European Young Statisticians Meeting (EYSM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.17089 2023-09-21 cs.LG cs.CL 57%

Concept-Oriented Deep Learning with Large Language Models

Daniel T. Chang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08932 2023-09-19 cs.CV 57%

Semantics-aware LiDAR-Only Pseudo Point Cloud Generation for 3D Object Detection

Tiago Cortinhal, Idriss Gouigah, Eren Erdal Aksoy

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03275 2023-09-19 cs.CV 57%

That's What I Said: Fully-Controllable Talking Face Generation

Youngjoon Jang, Kyeongha Rho, Jong-Bin Woo, Hyeongkeun Lee, Jihwan Park, Youshin Lim, Byeong-Yeol Kim, Joon Son Chung

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08888 2023-09-12 cs.CV 57%

Stochastic Segmentation with Conditional Categorical Diffusion Models

Lukas Zbinden, Lars Doorenbos, Theodoros Pissas, Adrian Thomas Huber, Raphael Sznitman, Pablo Márquez-Neila

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted at ICCV 2023. Code available at https://github.com/LarsDoorenbos/ccdm-stochastic-segmentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04287 2023-09-11 eess.SP cs.AI 57%

Sequential Semantic Generative Communication for Progressive Text-to-Image Generation

Hyelin Nam, Jihong Park, Jinho Choi, Seong-Lyun Kim

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments 4 pages, 2 figures, to be published in IEEE International Conference on Sensing, Communication, and Networking, Workshop on Semantic Communication for 6G (SC6G-SECON23)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00796 2023-09-06 cs.CV 57%

AttT2M: Text-Driven Human Motion Generation with Multi-Perspective Attention Mechanism

Chongyang Zhong, Lei Hu, Zihao Zhang, Shihong Xia

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments IEEE International Conference on Computer Vision 2023, 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.16758 2023-09-01 cs.CV 57%

Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only Images

Cuican Yu, Guansong Lu, Yihan Zeng, Jian Sun, Xiaodan Liang, Huibin Li, Zongben Xu, Songcen Xu, Wei Zhang, Hang Xu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.16632 2023-09-01 cs.CV 57%

3D-STMN: Dependency-Driven Superpoint-Text Matching Network for End-to-End 3D Referring Expression Segmentation

Changli Wu, Yiwei Ma, Qi Chen, Haowei Wang, Gen Luo, Jiayi Ji, Xiaoshuai Sun

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11510 2023-08-22 cs.CV 57%

Pushing the Limits of 3D Shape Generation at Scale

Yu Wang, Xuelin Qian, Jingyang Huo, Tiejun Huang, Bo Zhao, Yanwei Fu

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Project page: https://argus-3d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13233 2023-08-22 cs.CV 57%

Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open World

Qifan Yu, Juncheng Li, Yu Wu, Siliang Tang, Wei Ji, Yueting Zhuang

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.12249 2023-08-22 cs.CV cs.LG 57%

Do DALL-E and Flamingo Understand Each Other?

Hang Li, Jindong Gu, Rajat Koner, Sahand Sharifzadeh, Volker Tresp

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted to ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08157 2023-08-17 cs.CV 57%

Learning to Generate Semantic Layouts for Higher Text-Image Correspondence in Text-to-Image Synthesis

Minho Park, Jooyeol Yun, Seunghwan Choi, Jaegul Choo

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06739 2023-08-15 cs.CV 57%

Free-ATM: Exploring Unsupervised Learning on Diffusion-Generated Images with Free Attention Masks

David Junhao Zhang, Mutian Xu, Chuhui Xue, Wenqing Zhang, Xiaoguang Han, Song Bai, Mike Zheng Shou

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03678 2023-08-14 eess.IV cs.CV 57%

Towards Segment Anything Model (SAM) for Medical Image Segmentation: A Survey

Yichi Zhang, Rushi Jiao

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00906 2023-08-03 cs.CV 57%

ImageBrush: Learning Visual In-Context Instructions for Exemplar-Based Image Manipulation

Yasheng Sun, Yifan Yang, Houwen Peng, Yifei Shen, Yuqing Yang, Han Hu, Lili Qiu, Hideki Koike

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.16106 2023-08-01 cs.RO cs.CV 57%

TransFusion: A Practical and Effective Transformer-based Diffusion Model for 3D Human Motion Prediction

Sibo Tian, Minghui Zheng, Xiao Liang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.13908 2023-07-27 cs.CV 57%

Points-to-3D: Bridging the Gap between Sparse Points and Shape-Controllable Text-to-3D Generation

Chaohui Yu, Qiang Zhou, Jingliang Li, Zhe Zhang, Zhibin Wang, Fan Wang

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments Accepted by ACMMM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.05564 2023-07-13 cs.CL 57%

Augmenters at SemEval-2023 Task 1: Enhancing CLIP in Handling Compositionality and Ambiguity for Zero-Shot Visual WSD through Prompt Augmentation and Text-To-Image Diffusion

Jie S. Li, Yow-Ting Shiue, Yong-Siang Shih, Jonas Geiping

专题命中 多模态生成 :image-text(abstract);分类 cs.CL

Comments Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04978 2023-07-12 cs.CV 57%

Diffusion idea exploration for art generation

Nikhil Verma

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Report Submitted for degree completion of Master of Science in Applied Computing at University of Toronto

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11137 2023-06-21 eess.IV cs.CV 57%

Deep Learning Framework with Multi-Head Dilated Encoders for Enhanced Segmentation of Cervical Cancer on Multiparametric Magnetic Resonance Imaging

Reza Kalantar, Sebastian Curcean, Jessica M Winfield, Gigin Lin, Christina Messiou, Matthew D Blackledge, Dow-Mu Koh

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.10730 2023-06-21 cs.CV 57%

UniG3D: A Unified 3D Object Generation Dataset

Qinghong Sun, Yangguang Li, ZeXiang Liu, Xiaoshui Huang, Fenggang Liu, Xihui Liu, Wanli Ouyang, Jing Shao

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14671 2023-06-16 cs.CL 57%

A Survey of Diffusion Models in Natural Language Processing

Hao Zou, Zae Myung Kim, Dongyeop Kang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

Comments We changed the title of the paper due to a conflict with a previous paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04652 2023-06-09 cs.CV 57%

Language Adaptive Weight Generation for Multi-task Visual Grounding

Wei Su, Peihan Miao, Huanzhang Dou, Gaoang Wang, Liang Qiao, Zheyang Li, Xi Li

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00301 2023-06-07 cs.LG cs.CL 57%

CapText: Large Language Model-based Caption Generation From Image Context and Description

Shinjini Ghosh, Sagnik Anupam

专题命中 多模态生成 :image-text(abstract);分类 cs.CL

Comments Update 6/6/23: Fixed typographic error in abstract

详情

展开后加载摘要…

URL PDF HTML 收藏