arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86623 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70082 篇

2312.01152 2023-12-05 eess.IV cs.CV 86%

Ultra-Resolution Cascaded Diffusion Model for Gigapixel Image Synthesis in Histopathology

Sarah Cechnicka, Hadrien Reynaud, James Ball, Naomi Simmonds, Catherine Horsfield, Andrew Smith, Candice Roufosse, Bernhard Kainz

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(title);分类 cs.CV

Comments MedNeurIPS 2023 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16555 2023-11-29 cs.CV 86%

Enhancing Scene Text Detectors with Realistic Text Image Synthesis Using Diffusion Models

Ling Fu, Zijie Wu, Yingying Zhu, Yuliang Liu, Xiang Bai

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13384 2023-10-26 eess.IV cs.CV cs.LG 86%

DiffInfinite: Large Mask-Image Synthesis via Parallel Random Patch Diffusion in Histopathology

Marco Aversa, Gabriel Nobis, Miriam Hägele, Kai Standvoss, Mihaela Chirica, Roderick Murray-Smith, Ahmed Alaa, Lukas Ruff, Daniela Ivanova, Wojciech Samek, Frederick Klauschen, Bruno Sanguinetti, Luis Oala

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10326 2023-08-21 cs.CV cs.LG 86%

Unsupervised Out-of-Distribution Detection with Diffusion Inpainting

Zhenzhen Liu, Jin Peng Zhou, Yufan Wang, Kilian Q. Weinberger

专题命中 扩散模型 :diffusion(title,abstract);inpainting(title);分类 cs.CV

Comments ICML 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05424 2023-08-16 eess.IV cs.CV cs.LG 86%

Echo from noise: synthetic ultrasound image generation using diffusion models for real image segmentation

David Stojanovski, Uxio Hermida, Pablo Lamata, Arian Beqiri, Alberto Gomez

专题命中 扩散模型 :diffusion(title,abstract);image generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.11710 2023-08-02 cs.CV 86%

Controlled and Conditional Text to Image Generation with Diffusion Prior

Pranav Aggarwal, Hareesh Ravi, Naveen Marri, Sachin Kelkar, Fengbin Chen, Vinh Khuc, Midhun Harikumar, Ritiz Tambi, Sudharshan Reddy Kakumanu, Purvak Lapsiya, Alvin Ghouas, Sarah Saber, Malavika Ramprasad, Baldo Faieta, Ajinkya Kale

专题命中 扩散模型 :diffusion(title,abstract);image generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01148 2023-07-07 cs.CV eess.IV 86%

Investigating Data Memorization in 3D Latent Diffusion Models for Medical Image Synthesis

Salman Ul Hassan Dar, Arman Ghanaat, Jannik Kahmann, Isabelle Ayx, Theano Papavassiliu, Stefan O. Schoenberg, Sandy Engelhardt

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.01138 2023-05-03 eess.IV cs.CV 86%

High-Fidelity Image Synthesis from Pulmonary Nodule Lesion Maps using Semantic Diffusion Model

Xuan Zhao, Benjamin Hou

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(title);分类 cs.CV

Comments 4 pages, 1 figure, submitted to MIDL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02012 2023-04-14 cs.CV cs.AI cs.LG 86%

EGC: Image Generation and Classification via a Diffusion Energy-Based Model

Qiushan Guo, Chuofan Ma, Yi Jiang, Zehuan Yuan, Yizhou Yu, Ping Luo

专题命中 扩散模型 :image generation(title,abstract);diffusion(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.11477 2023-03-22 eess.IV cs.CV q-bio.QM 86%

NASDM: Nuclei-Aware Semantic Histopathology Image Generation Using Diffusion Models

Aman Shrivastava, P. Thomas Fletcher

专题命中 扩散模型 :diffusion(title,abstract);image generation(title);分类 cs.CV

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.08664 2023-03-17 cs.CV cs.LG 86%

Enhancing Diffusion-Based Image Synthesis with Robust Classifier Guidance

Bahjat Kawar, Roy Ganz, Michael Elad

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(title);分类 cs.CV

Comments Accepted to TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11138 2022-11-22 cs.CV 86%

Diffusion-Based Scene Graph to Image Generation with Masked Contrastive Pre-Training

Ling Yang, Zhilin Huang, Yang Song, Shenda Hong, Guohao Li, Wentao Zhang, Bin Cui, Bernard Ghanem, Ming-Hsuan Yang

专题命中 扩散模型 :diffusion(title,abstract);image generation(title);分类 cs.CV

Comments Code and models shall be released at https://github.com/YangLing0818/SGDiff

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.15282 2021-12-20 cs.CV cs.AI cs.LG 86%

Cascaded Diffusion Models for High Fidelity Image Generation

Jonathan Ho, Chitwan Saharia, William Chan, David J. Fleet, Mohammad Norouzi, Tim Salimans

专题命中 扩散模型 :diffusion(title,abstract);image generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.13606 2021-11-29 cs.LG cs.CV stat.ML 86%

Conditional Image Generation with Score-Based Diffusion Models

Georgios Batzolis, Jan Stanczuk, Carola-Bibiane Schönlieb, Christian Etmann

专题命中 扩散模型 :diffusion(title,abstract);image generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.12701 2021-11-25 cs.CV cs.LG 86%

Unleashing Transformers: Parallel Token Prediction with Discrete Absorbing Diffusion for Fast High-Resolution Image Generation from Vector-Quantized Codes

Sam Bond-Taylor, Peter Hessey, Hiroshi Sasaki, Toby P. Breckon, Chris G. Willcocks

专题命中 扩散模型 :diffusion(title,abstract);image generation(title);分类 cs.CV

Comments 19 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.10406 2020-06-19 cs.CV 86%

Fourth-Order Anisotropic Diffusion for Inpainting and Image Compression

Ikram Jumakulyyev, Thomas Schultz

专题命中 扩散模型 :diffusion(title,abstract);inpainting(title);分类 cs.CV

Comments Accepted for publication in Springer book "Anisotropy Across Fields and Scales"

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13247 2026-06-12 cs.AI 新提交 86%

EPIG: Emotion-Based Prompting for Personalised Image Generation

EPIG:基于情感提示的个性化图像生成

Emna Othmen, Mohamed Yassine Landolsi, Lotfi Ben Romdhane

机构 * MARS Research Lab LR17ES05, ISITCom, University of Sousse(苏塞大学ISITCom学院MARS研究实验室LR17ES05)

专题命中 扩散模型 :image generation(title,abstract);text-to-image(abstract,comments);diffusion(abstract,comments)

AI总结 提出EPIG方法,利用心理学效价-唤醒模型在提示层面增强情感表达,无需训练即可控制生成图像的唤醒度,在10个多样化提示上平均唤醒误差降低14%-17%。

Comments Submitted to arXiv. 20 pages, 4 figures. Work on emotion-based prompt engineering for text-to-image diffusion models with applications in personalized image generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16255 2026-04-06 astro-ph.IM cs.AI 86%

Category-based Galaxy Image Generation via Diffusion Models

基于类别的银河图像生成:通过扩散模型

Xingzhong Fan, Hongming Tang, Yue Zeng, M. B. N. Kouwenhoven, Guangquan Zeng

机构 * Department of Physics, Xi'an Jiaotong-Liverpool University(西交利物浦大学物理系) Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系) Department of Physics, The Chinese University of Hong Kong(香港中文大学物理系)

专题命中 扩散模型 :diffusion(title,abstract);image generation(title)

AI总结 本文提出GalCatDiff框架,结合银河图像特征和天体物理属性,通过增强的U-Net和新型Astro-RAB模块提升生成质量,实现高效且物理一致的银河生成。

Comments 23 pages, 10 figures. Accepted by AAS Astronomical Journal (AJ) and has now been published on https://iopscience.iop.org/article/10.3847/1538-3881/ae5064. See another independent work for further reference -- Can AI Dream of Unseen Galaxies? Conditional Diffusion Model for Galaxy Morphology Augmentation (Ma, Sun et al.). Comments are welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01134 2026-08-14 cs.GR cs.CV 版本更新 86%

RealMat: Realistic Materials with Diffusion and Reinforcement Learning

RealMat:结合扩散模型与强化学习的真实材质生成器

Xilong Zhou, Pedro Figueiredo, Miloš Hašan, Valentin Deschaintre, Paul Guerrero, Yiwei Hu, Nima Khademi Kalantari

机构 * Max Planck Institute for Informatics Saarbrücken Germany Texas A\&M University College Station USA Adobe Research San Jose USA Adobe Research London UK Max Planck Institute for Informatics Texas A\&M University Adobe Research

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 针对材质生成的合成数据真实感不足、真实数据规模有限的问题,提出结合SDXL微调与强化学习的RealMat,有效提升了生成材质的真实性。

Comments 12 pages, 12 figures

Journal ref Computer Graphics Forum 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19344 2026-07-22 cs.CV cs.AI cs.GR 新提交 86%

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

外观指针——扩散变压器的多模态区域控制

Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch, Srinath Sridhar, Matheus Gadelha

机构 * Brown University(布朗大学) Adobe Research(Adobe 研究院)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);image synthesis(abstract);分类 cs.CV、cs.GR

AI总结 针对可控图像生成难题,提出外观指针方法,通过区域对应网络和空间聚合机制生成并细化指针,为扩散变压器引入模态无关的局部多模态控制接口,单一模型性能达或超现有技术。

Comments 38 Pages, Preprint with supplement

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01929 2026-03-25 cs.GR cs.AI cs.CV cs.LG 86%

Image Generation from Contextually-Contradictory Prompts

从语境矛盾提示生成图像

Saar Huberman, Or Patashnik, Omer Dahary, Ron Mokady, Daniel Cohen-Or

机构 * Tel Aviv University(特拉维夫大学) BRIA AI(BRIA人工智能)

专题命中 扩散模型 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.GR

AI总结 本文提出一种阶段感知提示分解框架,通过代理提示引导去噪过程,解决提示中概念矛盾导致的语义不准确问题,提升图像生成的准确性。

Comments Project page: https://tdpc2025.github.io/SAP/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02580 2026-03-18 cs.CV cs.AI cs.GR cs.LG 86%

TAUE: Training-free Noise Transplant and Cultivation Diffusion Model

TAUE:无需训练的噪声移植与培育扩散模型

Daichi Nagai, Ryugo Morita, Shunsuke Kitada, Hitoshi Iyatomi

机构 * Faculty of Science and Engineering, Hosei University(恒河大学科学与工程学院) RPTU Kaiserslautern-Landau & DFKI GmbH(凯撒斯劳滕-兰道大学与DFKI GmbH)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 TAUE提出无需训练的噪声移植与培育扩散模型,通过嵌入全局结构信息和语义线索,实现多层图像生成,提升跨层一致性并支持新应用。

Comments Accepted to CVPR 2026 Findings. The first two authors contributed equally. Project Page: https://iyatomilab.github.io/TAUE

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05106 2026-03-06 cs.CV cs.GR cs.LG cs.RO 86%

NeuralRemaster: Phase-Preserving Diffusion for Structure-Aligned Generation

NeuralRemaster: 保留相位的扩散用于结构对齐生成

Yu Zeng, Charles Ochoa, Mingyuan Zhou, Vishal M. Patel, Vitor Guizilini, Rowan McAllister

机构 * Toyota Research Institute(丰田研究院) University of Texas, Austin(德克萨斯大学奥斯汀分校) Johns Hopkins University(约翰霍普金斯大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 NeuralRemaster通过保留相位并随机化幅度,实现结构对齐的图像和视频生成,提升模拟到现实的转换性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24108 2026-03-04 cs.GR cs.AI cs.CV cs.LG 86%

Navigating with Annealing Guidance Scale in Diffusion Space

在扩散空间中利用退火引导尺度导航

Shai Yehezkel, Omer Dahary, Andrey Voynov, Daniel Cohen-Or

机构 * Tel Aviv University(特拉维夫大学) Google DeepMind(谷歌DeepMind)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 本文提出了一种退火引导调度器,通过动态调整引导尺度提升文本到图像生成的质量和对齐度,无需额外资源消耗。

Comments SIGGRAPH Asia, 2025. Project page: https://annealing-guidance.github.io/annealing-guidance/

Journal ref ACM Trans. Graph., Vol. 44, No. 6, Article 5. Publication date: December 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17927 2026-01-27 cs.CV cs.MM 86%

RemEdit: Efficient Diffusion Editing with Riemannian Geometry

RemEdit: 基于黎曼几何的高效扩散编辑

Eashan Adhikarla, Brian D. Davison

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);image editing(abstract);分类 cs.CV、cs.MM

AI总结 RemEdit通过基于黎曼几何的潜在空间导航和任务特定注意力剪枝机制,实现了高效且保真的图像编辑,同时保持实时性能。

Journal ref IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17259 2026-01-27 cs.CV cs.GR cs.LG 86%

Inference-Time Loss-Guided Colour Preservation in Diffusion Sampling

推理时的损失引导颜色保持在扩散采样中

Angad Singh Ahuja, Aarush Ram Anandh

机构 * Constrained Image-Synthesis Lab(受限图像合成实验室)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);inpainting(abstract);分类 cs.CV、cs.GR

AI总结 本文提出一种无需额外训练的推理时颜色保持方法,通过区域约束和复合损失引导扩散模型,实现精准颜色控制。

Comments 25 Pages, 12 Figures, 3 Tables, 5 Appendices, 8 Algorithms

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05532 2025-10-08 cs.CV cs.GR cs.LG 86%

Teamwork: Collaborative Diffusion with Low-rank Coordination and Adaptation

Sam Sartor, Pieter Peers

机构 * College of William \& Mary Williamsburg USA College of William \& Mary

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11054 2025-07-03 cs.GR cs.CV cs.LG 86%

LUSD: Localized Update Score Distillation for Text-Guided Image Editing

Worameth Chinchuthakun, Tossaporn Saengja, Nontawat Tritrong, Pitchaporn Rewatbowornwong, Pramook Khungurn, Supasorn Suwajanakorn

专题命中 扩散模型 :image editing(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.GR

Comments ICCV 2025. Project page: https://github.com/sincostanx/LUSD

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12072 2024-11-20 cs.CV cs.AI cs.LG cs.MM 86%

Zoomed In, Diffused Out: Towards Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution

Brian B. Moser, Stanislav Frolov, Tobias C. Nauen, Federico Raue, Andreas Dengel

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23775 2024-11-06 cs.CV cs.GR 86%

In-Context LoRA for Diffusion Transformers

Lianghua Huang, Wei Wang, Zhi-Fan Wu, Yupeng Shi, Huanzhang Dou, Chen Liang, Yutong Feng, Yu Liu, Jingren Zhou

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.GR

Comments Tech report. Project page: https://ali-vilab.github.io/In-Context-LoRA-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏