arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86623 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70082 篇

2304.09479 2026-05-13 cs.CV cs.GR cs.LG 84%

DiFaReli++: Diffusion Face Relighting with Consistent Cast Shadows

DiFaReli++: 基于扩散的面部光照重建与一致阴影生成

Puntawat Ponglertnapakorn, Nontawat Tritrong, Supasorn Suwajanakorn

机构 * School of Information Science and Technology, Vidyasirimedhi Institute of Science and Technology(信息科学与技术学院,维达亚西里米迪科学技术研究所)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV、cs.GR

AI总结 本文提出一种单视图面部光照重建方法,通过条件扩散隐式模型实现无光照地面真值训练,实现真实光照下的阴影一致性。

Comments Published in IEEE TPAMI (vol. 48, no. 5, May 2026). This is an extended version of the ICCV 2023 paper (DiFaReli)

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 5, pp. 5068-5082, May 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00548 2026-05-12 cs.CV cs.GR 84%

Colorful-Noise: Training-Free Low-Frequency Noise Manipulation for Color-Based Conditional Image Generation

彩色噪声:无需训练的低频噪声操控用于基于颜色的条件图像生成

Nadav Z. Cohen, Ofir Abramovich, Ariel Shamir

机构 * Reichman University(雷曼大学)

专题命中 扩散模型 :image generation(title);text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.GR

AI总结 本文研究了扩散模型输入噪声的特性,发现低频成分主导图像全局结构和颜色,高频频成分控制细节。通过低频图像先验操控低频噪声,实现无需训练的条件生成,控制整体结构和颜色,保留高频细节多样性。

Comments SIGGRAPH 2026 Conference Paper. Project Page at: https://nadavc220.github.io/colorful-noise/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01540 2026-04-17 cs.CV cs.AI cs.GR cs.LG 84%

Edge-preserving noise for diffusion models

边缘保留噪声用于扩散模型

Jente Vandersanden, Sascha Holl, Xingchang Huang, Gurprit Singh

机构 * Max Planck Institute for Informatics(马克斯·普朗克信息研究所) Advanced Micro Devices(先进微器件)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.GR

AI总结 本文提出一种边缘保留的扩散过程,通过混合噪声方案在边缘感知调度器中平滑过渡,提升结构细节捕捉能力,同时保持全局性能,并在图像生成和结构引导任务中取得改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06161 2026-04-13 cs.CV cs.AI cs.GR 84%

DiffHDR: Re-Exposing LDR Videos with Video Diffusion Models

DiffHDR:利用视频扩散模型重新暴露LDR视频

Zhengming Yu, Li Ma, Mingming He, Leo Isikdogan, Yuancheng Xu, Dmitriy Smirnov, Pablo Salamanca, Dao Mi, Pablo Delgado, Ning Yu, Julien Philip, Xin Li, Wenping Wang, Paul Debevec

机构 * Eyeline Labs(Eyeline 实验室) Netflix

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

AI总结 DiffHDR通过视频扩散模型的潜在空间将LDR视频转换为HDR,恢复过曝和欠曝区域的真实细节,提升动态范围和再曝光效果。

Comments 28 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07455 2026-03-31 cs.CV cs.AI cs.CL cs.GR 84%

Image Generation Models: A Technical History

图像生成模型:技术史

Rouzbeh Shirvani

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV、cs.GR

AI总结 本文综述了过去十年图像生成模型的快速发展,涵盖VAEs、GANs、正常化流、自回归和Transformer模型以及扩散方法,分析其技术原理、训练方法及局限性,并讨论视频生成和模型鲁棒性等最新进展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13739 2026-03-17 cs.CV cs.AI cs.MM 84%

UniVid: Pyramid Diffusion Model for High Quality Video Generation

UniVid:金字塔扩散模型用于高质量视频生成

Xinyu Xiao, Binbin Yang, Tingtian Li, Yipeng Yu, Sen Lei

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.MM

AI总结 UniVid通过融合文本提示和参考图像,结合时间金字塔交叉帧空间时间注意力模块,提升文本到视频、图像到视频及联合生成任务的时序一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09411 2026-02-10 cs.CV cs.GR 84%

ImageRAG: Dynamic Image Retrieval for Reference-Guided Image Generation

ImageRAG:用于参考引导图像生成的动态图像检索

Rotem Shalev-Arkushin, Rinon Gal, Amit H. Bermano, Ohad Fried

机构 * Tel-Aviv University(特拉维夫大学) Tel Aviv University(特拉维夫大学) Reichman University(里奇曼大学)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV、cs.GR

AI总结 ImageRAG通过动态图像检索提升参考引导图像生成的质量,无需专门训练即可适应不同模型类型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09131 2026-02-04 cs.GR cs.AI cs.CV 84%

Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer

无需训练的文本引导颜色编辑与多模态扩散变换器

Zixin Yin, Xili Dai, Ling-Hao Chen, Deyu Zhou, Jianan Wang, Duomin Wang, Gang Yu, Lionel M. Ni, Lei Zhang, Heung-Yeung Shum

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) International Digital Economy Academy(国际数字经济学院) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Tsinghua University(清华大学) Astribot StepFun

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV、cs.GR

AI总结 ColorCtrl通过多模态扩散变换器实现无需训练的文本引导颜色编辑,精准控制颜色属性并保持一致性,优于现有方法和商业模型。

Comments https://zxyin.github.io/ColorCtrl

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15468 2026-01-13 cs.CV cs.GR 84%

Semantic Aware Diffusion Inverse Tone Mapping

语义感知的扩散反色调映射

Abhishek Goswami, Aru Ranjan Singh, Francesco Banterle, Kurt Debattista, Thomas Bashford-Rogers

机构 * University of Warwick(沃里克大学)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

AI总结 本文提出一种基于语义感知扩散的反色调映射方法,通过生成被剪裁区域的丢失细节,提升SDR图像至HDR,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23484 2025-12-23 cs.MM cs.CV eess.IV 84%

TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion Sensitivity

TAG-WM: 通过扩散反向敏感性实现的抗篡改生成图像水印

Yuzhuo Chen, Zehua Ma, Han Fang, Weiming Zhang, Nenghai Yu

机构 * Anhui Province Key Laboratory of Digital Security, University of Science and Technology of China(安徽省数字安全重点实验室,中国科学技术大学) School of Computing, National University of Singapore(computing学院,新加坡国立大学)

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV、cs.MM

AI总结 TAG-WM通过扩散反向敏感性实现抗篡改生成图像水印,提升篡改鲁棒性和定位能力,保持生成质量与水印容量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16670 2025-12-19 cs.CV cs.GR 84%

FrameDiffuser: G-Buffer-Conditioned Diffusion for Neural Forward Frame Rendering

FrameDiffuser: 基于G-Buffer的扩散神经渲染

Ole Beisswenger, Jan-Niklas Dihlmann, Hendrik P. A. Lensch

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.GR

AI总结 FrameDiffuser通过基于G-buffer的自回归框架实现时间一致的神经渲染,结合结构引导和时间一致性,提升逼真度和效率。

Comments Project Page: https://framediffuser.jdihlmann.com/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09441 2025-12-12 cs.GR cs.CV 84%

RectifiedHR: High-Resolution Diffusion via Energy Profiling and Adaptive Guidance Scheduling

RectifiedHR: 通过能量分析和自适应引导调度实现高分辨率扩散

Ankit Sanjyal

机构 * Fordham University(福特汉姆大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.GR

AI总结 RectifiedHR通过能量分析和自适应引导调度提升高分辨率扩散模型的稳定性和图像质量。

Comments 8 Pages, 10 Figures, Pre-Print Version, This version is under review for citation accuracy

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05198 2025-12-08 cs.CV cs.GR cs.LG 84%

Your Latent Mask is Wrong: Pixel-Equivalent Latent Compositing for Diffusion Models

你的潜在掩码是错的:用于扩散模型的像素等效潜在合成

Rowan Bradbury, Dazhi Zhong

机构 * Bradbury Group(布拉德伯格集团)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

AI总结 本文提出PELC原则,通过DecFormer实现像素等效潜在合成,提升扩散模型在补全任务中的性能与保真度。

Comments 16 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08530 2025-10-10 cs.GR cs.CV 84%

X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering

Zhitong Huang, Mohan Zhang, Renhan Wang, Rui Tang, Hao Zhu, Jing Liao

机构 * City University of Hong Kong(香港城市大学) WeChat, Tencent Inc(微信、腾讯公司) Manycore Tech Inc(很多核科技公司)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV、cs.GR

Comments Code, model, and dataset will be released at project page soon: https://luckyhzt.github.io/x2video

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24369 2025-09-30 cs.CV cs.AI cs.MM 84%

From Satellite to Street: A Hybrid Framework Integrating Stable Diffusion and PanoGAN for Consistent Cross-View Synthesis

Khawlah Bajbaa, Abbas Anwar, Muhammad Saqib, Hafeez Anwar, Nabin Sharma, Muhammad Usman

机构 * Department of Information and Computer Science, King Fahd University of Petroleum and Minerals(信息与计算机科学系,国王法赫德石油与矿物大学) NCMI, CSIRO(CSIRO国家科学机构) Department of Computer Science, National University of Computer and Emerging Sciences (FAST-NUCES)(计算机科学系,国家计算机与新兴科学大学(FAST-NUCES)) Faculty of Engineering and IT, University of Technology Sydney(工程与信息技术学院,技术大学悉尼) Faculty of Science, University of Ontario Institute of Technology(科学学院, Ontario Institute of Technology大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21887 2025-09-29 cs.CV cs.MM 84%

StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing

Liyang Chen, Tianze Zhou, Xu He, Boshi Tang, Zhiyong Wu, Yang Huang, Yang Wu, Zhongqian Sun, Wei Yang, Helen Meng

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03680 2025-09-05 cs.GR cs.AI cs.CV 84%

LuxDiT: Lighting Estimation with Video Diffusion Transformer

Ruofan Liang, Kai He, Zan Gojcic, Igor Gilitschenski, Sanja Fidler, Nandita Vijaykumar, Zian Wang

机构 * NVIDIA University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.GR

Comments Project page: https://research.nvidia.com/labs/toronto-ai/LuxDiT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16827 2025-08-26 cs.GR cs.CV cs.LG 84%

Beyond Blur: A Fluid Perspective on Generative Diffusion Models

Grzegorz Gruszczynski, Jakub Meixner, Michal Jan Wlodarczyk, Przemyslaw Musialski

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.GR

Comments ICCV 2025 main conference, 8 pages paper, 20 pages appendix, 24 figures, supplementary pseudocode in appendix, https://iccv.thecvf.com/virtual/2025/poster/1176

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08384 2025-08-13 cs.GR cs.AI cs.CV 84%

Spatiotemporally Consistent Indoor Lighting Estimation with Diffusion Priors

Mutian Tong, Rundi Wu, Changxi Zheng

机构 * Columbia University(哥伦比亚大学)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

Comments 11 pages. Accepted by SIGGRAPH 2025 as Conference Paper

Journal ref SIGGRAPH '25: ACM SIGGRAPH 2025 Conference Conference Papers, Article 107, pages1-11, July 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19853 2025-08-05 cs.CV cs.GR cs.LG 84%

Conditional Balance: Improving Multi-Conditioning Trade-Offs in Image Generation

Nadav Z. Cohen, Oron Nir, Ariel Shamir

机构 * Reichman University(里奇曼大学) Microsoft Corporation(微软公司)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV、cs.GR

Comments Conference paper at CVPR 2025. Project page: https://nadavc220.github.io/conditional-balance.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15399 2025-07-22 cs.GR cs.CV 84%

Blended Point Cloud Diffusion for Localized Text-guided Shape Editing

Etai Sella, Noam Atia, Ron Mokady, Hadar Averbuch-Elor

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

Comments Accepted to ICCV 2025. Project Page: https://tau-vailab.github.io/BlendedPC/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16977 2025-05-23 cs.CV cs.MM 84%

Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On

Siqi Wan, Jingwen Chen, Yingwei Pan, Ting Yao, Tao Mei

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV、cs.MM

Comments ICLR 2025. Code is publicly available at: https://github.com/HiDream-ai/SPM-Diff

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10558 2025-05-16 cs.GR cs.CV 84%

Style Customization of Text-to-Vector Generation with Image Diffusion Priors

Peiying Zhang, Nanxuan Zhao, Jing Liao

机构 * City University of Hong Kong(香港城市大学) Adobe Research(Adobe研究)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR

Comments Accepted by SIGGRAPH 2025 (Conference Paper). Project page: https://customsvg.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16938 2025-04-14 cs.CV cs.AI cs.GR 84%

Generative Object Insertion in Gaussian Splatting with a Multi-View Diffusion Model

Hongliang Zhong, Can Wang, Jingbo Zhang, Jing Liao

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

Comments Accepted by Visual Informatics. Project Page: https://github.com/JiuTongBro/MultiView_Inpaint

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21694 2025-03-28 cs.GR cs.AI cs.CV 84%

Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D Data

Zhiyuan Ma, Xinyue Liang, Rongyuan Wu, Xiangyu Zhu, Zhen Lei, Lei Zhang

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR

Comments Accepted to CVPR 2025. Code:https://github.com/theEricMa/TriplaneTurbo. Demo:https://huggingface.co/spaces/ZhiyuanthePony/TriplaneTurbo

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18590 2025-03-25 cs.CV cs.GR 84%

DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models

Ruofan Liang, Zan Gojcic, Huan Ling, Jacob Munkberg, Jon Hasselgren, Zhi-Hao Lin, Jun Gao, Alexander Keller, Nandita Vijaykumar, Sanja Fidler, Zian Wang

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV、cs.GR

Comments CVPR 2025; project page: research.nvidia.com/labs/toronto-ai/DiffusionRenderer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04776 2025-03-10 cs.GR cond-mat.mtrl-sci cs.CV cs.LG 84%

GrainPaint: A multi-scale diffusion-based generative model for microstructure reconstruction of large-scale objects

Nathan Hoffman, Cashen Diniz, Dehao Liu, Theron Rodgers, Anh Tran, Mark Fuge

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04378 2025-02-10 cs.CV cs.GR cs.LG cs.SE 84%

DILLEMA: Diffusion and Large Language Models for Multi-Modal Augmentation

Luciano Baresi, Davide Yi Xian Hu, Muhammad Irfan Mas'udi, Giovanni Quattrocchi

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02225 2025-02-05 cs.CV cs.AI cs.MM 84%

Exploring the latent space of diffusion models directly through singular value decomposition

Li Wang, Boyan Gao, Yanran Li, Zhao Wang, Xiaosong Yang, David A. Clifton, Jun Xiao

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18390 2024-12-30 cs.CV cs.AI cs.LG cs.MM 84%

RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction

Xiaoping Wu, Jie Hu, Xiaoming Wei

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.MM

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏