arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86714 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3486 篇

2301.09637 2023-08-16 cs.CV cs.AI cs.GR cs.LG 62%

InfiniCity: Infinite-Scale City Synthesis

Chieh Hubert Lin, Hsin-Ying Lee, Willi Menapace, Menglei Chai, Aliaksandr Siarohin, Ming-Hsuan Yang, Sergey Tulyakov

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04343 2023-08-09 cs.CV cs.IR cs.MM 62%

Unifying Two-Stream Encoders with Transformers for Cross-Modal Retrieval

Yi Bin, Haoxuan Li, Yahui Xu, Xing Xu, Yang Yang, Heng Tao Shen

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

Comments Accepted at ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00148 2023-08-02 cs.CV cs.GR 62%

Controlling Geometric Abstraction and Texture for Artistic Images

Martin Büßemeyer, Max Reimann, Benito Buchheim, Amir Semmo, Jürgen Döllner, Matthias Trapp

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.04399 2023-04-11 cs.CV cs.AI cs.LG cs.MM 62%

CAVL: Learning Contrastive and Adaptive Representations of Vision and Language

Shentong Mo, Jingfei Xia, Ihor Markevych

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04868 2023-02-10 cs.CV cs.GR 62%

MEGANE: Morphable Eyeglass and Avatar Network

Junxuan Li, Shunsuke Saito, Tomas Simon, Stephen Lombardi, Hongdong Li, Jason Saragih

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Project page: https://junxuan-li.github.io/megane/

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.01806 2022-07-26 cs.CV cs.GR 62%

Neural Scene Decoration from a Single Photograph

Hong-Wing Pang, Yingshu Chen, Phuoc-Hieu Le, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments ECCV 2022 paper. 14 pages of main content, 4 pages of references, and 11 pages of appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.02308 2022-07-25 cs.CV cs.GR 62%

MoFaNeRF: Morphable Facial Neural Radiance Field

Yiyu Zhuang, Hao Zhu, Xusen Sun, Xun Cao

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments accepted to ECCV2022; code available at http://github.com/zhuhao-nju/mofanerf

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.02546 2022-07-05 cs.CV cs.GR 62%

GANSpace: Discovering Interpretable GAN Controls

Erik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain Paris

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Accepted to NeurIPS 2020

Journal ref Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 9841-9850

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.04455 2022-04-12 cs.GR cs.CV eess.IV 62%

Noise-based Enhancement for Foveated Rendering

Taimoor Tariq, Cara Tursun, Piotr Didyk

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments 14 pages including refences

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.05349 2022-03-11 cs.MM cs.CV 62%

Two-stream Hierarchical Similarity Reasoning for Image-text Matching

Ran Chen, Hanli Wang, Lei Wang, Sam Kwong

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.14020 2022-03-01 cs.CV cs.GR cs.LG 62%

State-of-the-Art in the Architecture, Methods and Applications of StyleGAN

Amit H. Bermano, Rinon Gal, Yuval Alaluf, Ron Mokady, Yotam Nitzan, Omer Tov, Or Patashnik, Daniel Cohen-Or

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13162 2022-03-01 cs.CV cs.AI cs.GR cs.LG 62%

Pix2NeRF: Unsupervised Conditional $π$-GAN for Single Image to Neural Radiance Fields Translation

Shengqu Cai, Anton Obukhov, Dengxin Dai, Luc Van Gool

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments 16 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.09855 2021-12-02 cs.CV cs.GR 62%

Infinite Nature: Perpetual View Generation of Natural Scenes from a Single Image

Andrew Liu, Richard Tucker, Varun Jampani, Ameesh Makadia, Noah Snavely, Angjoo Kanazawa

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments ICCV 2021 (oral); Project page: https://infinite-nature.github.io/; Video: https://www.youtube.com/watch?v=oXUf6anNAtc

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06688 2021-10-14 cs.CV cs.GR 62%

DeepVecFont: Synthesizing High-quality Vector Fonts via Dual-modality Learning

Yizhi Wang, Zhouhui Lian

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments SIGGRAPH Asia 2021 Technical Paper. Code: https://github.com/yizhiwang96/deepvecfont ; Homepage: https://yizhiwang96.github.io/deepvecfont_homepage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04654 2021-09-13 cs.GR cs.CV 62%

Per Garment Capture and Synthesis for Real-time Virtual Try-on

Toby Chong, I-Chao Shen, Nobuyuki Umetani, Takeo Igarashi

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Accepted to UIST2021. Project page: https://sites.google.com/view/deepmannequin/home

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.16011 2021-03-30 cs.CV cs.GR 62%

Intrinsic Autoencoders for Joint Neural Rendering and Intrinsic Image Decomposition

Hassan Abu Alhaija, Siva Karthik Mustikovela, Justus Thies, Varun Jampani, Matthias Nießner, Andreas Geiger, Carsten Rother

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.12884 2020-12-24 cs.CV cs.GR 62%

Vid2Actor: Free-viewpoint Animatable Person Synthesis from Video in the Wild

Chung-Yi Weng, Brian Curless, Ira Kemelmacher-Shlizerman

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Project Page: https://grail.cs.washington.edu/projects/vid2actor/ Supplementary Video: https://youtu.be/Zec8Us0v23o

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.04639 2020-08-11 cs.CV cs.GR 62%

Visual Indeterminacy in GAN Art

Aaron Hertzmann

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Leonardo / SIGGRAPH 2020 Art Papers

Journal ref Leonardo, Volume 53, Issue 4, August 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.02867 2020-04-07 cs.CV cs.GR 62%

Rethinking Spatially-Adaptive Normalization

Zhentao Tan, Dongdong Chen, Qi Chu, Menglei Chai, Jing Liao, Mingming He, Lu Yuan, Nenghai Yu

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.11378 2019-11-27 cs.LG cs.CV cs.MM eess.IV stat.ML 62%

Text2FaceGAN: Face Generation from Fine Grained Textual Descriptions

Osaid Rehman Nasir, Shailesh Kumar Jha, Manraj Singh Grover, Yi Yu, Ajit Kumar, Rajiv Ratn Shah

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.02714 2019-08-12 cs.GR cs.CV 62%

Relighting Humans: Occlusion-Aware Inverse Rendering for Full-Body Human Images

Yoshihiro Kanamori, Yuki Endo

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Published at SIGGRAPH Asia 2018 (ACM Transactions on Graphics). Project page with codes, pretrained models, and human model lists is at http://kanamori.cs.tsukuba.ac.jp/projects/relighting_human/

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.11478 2019-06-28 cs.CV cs.GR 62%

A Convolutional Decoder for Point Clouds using Adaptive Instance Normalization

Isaak Lim, Moritz Ibing, Leif Kobbelt

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Symposium on Geometry Processing 2019

Journal ref Computer Graphics Forum 38 (5), 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.12620 2019-04-30 cs.CV cs.CR cs.IR cs.MM 62%

AnonymousNet: Natural Face De-Identification with Measurable Privacy

Tao Li, Lei Lin

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.MM

Comments CVPR-19 Workshop on Computer Vision: Challenges and Opportunities for Privacy and Security (CV-COPS 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.06601 2018-12-04 cs.CV cs.GR cs.LG 62%

Video-to-Video Synthesis

Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, Bryan Catanzaro

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments In NeurIPS, 2018. Code, models, and more results are available at https://github.com/NVIDIA/vid2vid

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.04139 2023-05-05 cs.CV cs.AI 61%

DALLE-URBAN: Capturing the urban design expertise of large text to image transformers

Sachith Seneviratne, Damith Senanayake, Sanka Rasnayaka, Rajith Vidanaarachchi, Jason Thompson

专题命中 文生图 :text-to-image(abstract);分类 cs.CV;diffusion(comments)

Comments Accepted to DICTA 2022, released 11000+ environmental scene images generated by Stable Diffusion and 1000+ images generated by DALLE-2

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08724 2026-08-25 cs.CV 版本更新 57%

SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

SynerMedGen:通过任务对齐协同医疗多模态理解与生成

Weiren Zhao, Yi Dong, Cheng Chen

机构 * The University of Hong Kong, Hong Kong, China(香港大学)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

AI总结 SynerMedGen通过任务对齐实现医疗多模态理解与生成的协同,提出生成对齐理解任务和两阶段训练策略,实现零样本性能和跨数据集泛化,释放大规模数据集支持进一步研究。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20886 2026-08-24 cs.CV cs.LG 新提交 57%

EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

EviRank:用于多模态图像重排序的结构化相关性证据

Enjun Du, Siyi Liu, Zirong Chen, Xinyu Zuo, Jinwen Luo, Ruiwen Tao, Lisheng Duan, Haijin Liang, Jin Ma, Junfu Pu, Yongqi Zhang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tencent Yuanbao(腾讯元宝) The University of Hong Kong(香港大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 针对现有多模态图像重排序器的不足,提出EviRank将查询解析为结构化证据包,通过证据条件验证实现重排序,在五个基准上达SOTA,蒸馏学生模型保留超90%能力且成本更低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17564 2026-08-19 cs.CV cs.AI 新提交 57%

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

新概念必须进入之处:统一多模态模型中的入口门跨任务可用性

Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Columbia University(哥伦比亚大学) CUHK(香港中文大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 该研究通过分离统一多模态模型的理解与生成任务方向,发现跨任务可用性取决于概念绑定的入口层,提出的对齐目标可在极低损失下实现概念跨任务迁移。

Comments 27 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15395 2026-08-18 cs.CV 新提交 57%

JoLT: Joint Latent Trajectories for Context-Guided High-Resolution Tiled Generation

JoLT:用于上下文引导的高分辨率分块生成的联合潜在轨迹

Mathis Koroglu, Guillaume Jeanneret, Hugo Caselles-Dupré, Matthieu Cord, Arnaud Dapogny

机构 * Obvious Research(奥布弗西斯研究公司) Sorbonne Université(索邦大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 本文提出JoLT方法,通过联合去噪LR与HR潜在图像生成高分辨率图像,其生成的图像细节丰富、视觉效果佳,优于竞争基线,为艺术创作提供新方向。

Comments 25 pages, 10 figures, 7 tables. Accepted at the AI4VA Workshop at ECCV 2026. Project page: https://obvious-research.github.io/jolt/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00410 2026-08-18 cs.AI cs.CL cs.CV 版本更新 57%

Where did the ambiguity go? Examining how multimodal models interpret polysemous words

歧义去了哪里?探究多模态模型如何解释多义词

Jasin Cekinmez, Addison J. Wu, Raja Marjieh, Thomas L. Griffiths

机构 * Princeton University(普林斯顿大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

AI总结 该研究对比17个文本到图像模型和15个文本生成模型,发现多模态模型生成图像的词义多样性低于文本,揭示了基础模型在不同模态间意义表达的迁移 gap。

Comments Oral Presentation, Sci-FM Workshop @ COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏