arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 70159 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70159 篇

2512.00369 2025-12-02 cs.CV 83%

POLARIS: Projection-Orthogonal Least Squares for Robust and Adaptive Inversion in Diffusion Models

POLARIS:用于扩散模型中鲁棒和自适应逆向的投影正交最小二乘法

Wenshuo Chen, Haosen Li, Shaofeng Liang, Lei Wang, Haozhe Jia, Kaishen Yuan, Jieming Wu, Bowen Tian, Yutao Yue

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Griffith University(格里菲斯大学) Data61/CSIRO Project(Data61/CSIRO项目)

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV

AI总结 POLARIS通过将逆向问题转化为误差来源问题,以单行代码提升扩散模型逆向的鲁棒性和适应性,显著减少噪声近似误差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00084 2025-12-02 cs.CV cs.LG 83%

A Fast and Efficient Modern BERT based Text-Conditioned Diffusion Model for Medical Image Segmentation

一种快速且高效的基于现代BERT的文本条件扩散模型用于医学图像分割

Venkata Siddharth Dhara, Pawan Kumar

机构 * International Institute of Information Technology, Hyderabad, 500032, India(国际信息科技学院,海得拉巴)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文提出FastTextDiff,利用ModernBERT提升医学图像分割的效率和准确性,通过整合文本注释和多模态注意力机制改进传统扩散模型。

Comments 15 pages, 3 figures, Accepted in Slide 3 10th International Conference on Computer Vision & Image Processing (CVIP 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23377 2025-12-01 cs.CV 83%

DEAL-300K: Diffusion-based Editing Area Localization with a 300K-Scale Dataset and Frequency-Prompted Baseline

DEAL-300K:基于扩散的编辑区域定位与30万规模数据集及频率提示基线

Rui Zhang, Hongxia Wang, Hangqing Liu, Yang Zhou, Qiang Zeng

机构 * School of Cyber Science and Engineering, Sichuan University(四川大学计算机科学与工程学院) Key Laboratory of Data Protection and Intelligent Management, Ministry of Education, Sichuan University(教育部数据保护与智能管理重点实验室)

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV

AI总结 DEAL-300K通过多模态大语言模型生成编辑指令,结合频率提示微调方法,实现了基于扩散的图像编辑区域定位,提供高精度的基线和研究基础。

Comments 13pages,12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19990 2025-11-26 cs.CV 83%

OmniRefiner: Reinforcement-Guided Local Diffusion Refinement

OmniRefiner: 基于强化学习的局部扩散细化

Yaoli Liu, Ziheng Ouyang, Shengtao Lou, Yiren Song

机构 * Zhejiang University(浙江大学) Nankai University(南开大学) National University of Singapore(新加坡国立大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 OmniRefiner通过强化学习和局部扩散细化技术,提升参考引导图像生成中细粒度细节的保留与一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19936 2025-11-26 cs.CV 83%

Image Diffusion Models Exhibit Emergent Temporal Propagation in Videos

图像扩散模型在视频中表现出涌现的时间传播

Youngseo Kim, Dohyun Kim, Geonhee Han, Paul Hongsuck Seo

机构 * Dept. of CSE, Korea University(计算机科学与工程系,韩国大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文提出DRIFT框架,利用预训练图像扩散模型和SAM引导的掩码细化,实现视频中鲁棒的零样本物体跟踪。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19910 2025-11-26 eess.IV cs.CV 83%

DLADiff: A Dual-Layer Defense Framework against Fine-Tuning and Zero-Shot Customization of Diffusion Models

DLADiff: 一种双层防御框架,用于对抗扩散模型的微调和零样本定制

Jun Jia, Hongyi Miao, Yingjie Zhou, Linhan Cao, Yanwei Jiang, Wangqiu Zhou, Dandan Zhu, Hua Yang, Wei Sun, Xiongkuo Min, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Shandong University(山东大学) Hefei University of Technology(合肥工业大学) East China Normal University(华东师范大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 DLADiff提出一种双层防御框架,通过双代理模型和交替动态微调有效防御扩散模型的微调和零样本生成攻击。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19385 2025-11-26 cs.CV cs.AI 83%

Advancing Limited-Angle CT Reconstruction Through Diffusion-Based Sinogram Completion

通过扩散基于的sinogram补全推进有限角度CT重建

Jiaqi Guo, Santiago Lopez-Tapia, Aggelos K. Katsaggelos

机构 * Dept. of Electrical and Computer Engineering, Northwestern University, Evanston, IL, USA(电气与计算机工程系,西北大学,爱荷华州埃文斯顿)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

AI总结 本文提出基于扩散模型的sinogram补全方法,结合蒸馏和伪逆约束,实现高效准确的有限角度CT重建,有效抑制伪影并保持结构细节。

Comments Accepted at the 2025 IEEE International Conference on Image Processing (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19434 2025-11-25 cs.CV cs.LG stat.ML 83%

Breaking the Likelihood-Quality Trade-off in Diffusion Models by Merging Pretrained Experts

通过合并预训练专家打破扩散模型中似然-质量权衡

Yasin Esfandiari, Stefan Bauer, Sebastian U. Stich, Andrea Dittadi

机构 * Saarland University(萨尔兰大学) Helmholtz AI(亥姆霍兹人工智能研究所) Technical University of Munich(慕尼黑技术大学) CISPA Helmholtz Center for Information Security(亥姆霍兹信息安全部分研究所) MPI for Intelligent Systems, Tübingen(图宾根智能系统研究所)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 通过在去噪过程中切换预训练专家,该方法有效打破扩散模型中似然与质量的权衡,提升图像生成质量和似然

Comments ICLR 2025 DeLTa workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01421 2025-11-25 cs.CV 83%

InfoScale: Unleashing Training-free Variable-scaled Image Generation via Effective Utilization of Information

InfoScale: 通过有效利用信息实现免训练可变尺度图像生成

Guohui Zhang, Jiangtong Tan, Linjiang Huang, Zhonghang Yuan, Mingde Yao, Jie Huang, Feng Zhao

机构 * USTC(中国科学技术大学) Beihang University(北京航空航天大学) CUHK MMLab(香港中文大学多媒体实验室) Kuaishou Technology(快手科技)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 InfoScale通过有效利用信息解决扩散模型在可变尺度图像生成中的信息丢失、聚合不灵活和分布不匹配问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16904 2025-11-24 cs.CV 83%

Warm Diffusion: Recipe for Blur-Noise Mixture Diffusion Models

Warm Diffusion: Blur-Noise Mixture Diffusion Models的配方

Hao-Chien Hsueh, Chi-En Yen, Wen-Hsiao Peng, Ching-Chun Huang

机构 * Department of Computer Science, National Yang Ming Chiao Tung University, Hsinchu, Taiwan(计算机科学系,National Yang Ming Chiao Tung大学,Hsinchu,台湾)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 Warm Diffusion通过融合模糊与噪声的混合模型,解决传统热扩散和冷扩散的不足,提升图像生成效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10629 2025-11-24 cs.CV 83%

One Small Step in Latent, One Giant Leap for Pixels: Fast Latent Upscale Adapter for Your Diffusion Models

潜在空间中的一小步,像素世界中的一大跃:一种快速的潜在上采样适配器用于您的扩散模型

Aleksandr Razin, Danil Kazantsev, Ilya Makarov

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 提出一种轻量级潜在上采样适配器,通过潜在空间单次前向传递实现高分辨率图像合成,提升效率并保持保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16317 2025-11-21 cs.CV 83%

NaTex: Seamless Texture Generation as Latent Color Diffusion

NaTex: 无缝纹理生成作为潜在颜色扩散

Zeqiang Lai, Yunfei Zhao, Zibo Zhao, Xin Yang, Xin Huang, Jingwei Huang, Xiangyu Yue, Chunchao Guo

机构 * MMLab, CUHK(CUHK多媒体实验室) Tencent Hunyuan(腾讯文生)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

AI总结 NaTex通过直接在3D空间中预测纹理颜色,提出潜在颜色扩散方法,实现高效且一致的纹理生成与重建。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14113 2025-11-19 cs.CV 83%

Coffee: Controllable Diffusion Fine-tuning

Ziyao Zeng, Jingcheng Ni, Ruyi Liu, Alex Wong

机构 * Yale University(耶鲁大学) Brown University(布朗大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13689 2025-11-19 cs.CL cs.CV 83%

Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation

Sofia Jamil, Kotla Sai Charan, Sriparna Saha, Koustava Goswami, Joseph K J

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09611 2025-11-19 cs.CV 83%

MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation

Ye Tian, Ling Yang, Jiongfan Yang, Anran Wang, Yu Tian, Jiani Zheng, Haochen Wang, Zhiyang Teng, Zhuochen Wang, Yinjie Wang, Yunhai Tong, Mengdi Wang, Xiangtai Li

机构 * Peking University(北京大学) ByteDance(字节跳动) Princeton University(普林斯顿大学) CASIA(中国科学院自动化研究所) The University of Chicago(芝加哥大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

Comments Project Page: https://tyfeld.github.io/mmadaparellel.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05772 2025-11-18 cs.CV 83%

MAISI-v2: Accelerated 3D High-Resolution Medical Image Synthesis with Rectified Flow and Region-specific Contrastive Loss

Can Zhao, Pengfei Guo, Dong Yang, Yucheng Tang, Yufan He, Benjamin Simon, Mason Belue, Stephanie Harmon, Baris Turkbey, Daguang Xu

专题命中 扩散模型 :image synthesis(title,abstract);diffusion(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06157 2025-11-18 cs.CV 83%

3D-free meets 3D priors: Novel View Synthesis from a Single Image with Pretrained Diffusion Guidance

Taewon Kang, Divya Kothandaraman, Dinesh Manocha, Ming C. Lin

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Accepted to The 40th Annual AAAI Conference on Artificial Intelligence (AAAI-26), AAAI 2026 Workshop on AI for Environmental Science (AI4ES). Due to arXiv's 1,920-character limit, the abstract here is shortened. Please refer to the paper (View PDF) to read the full abstract. 14 pages, 13 figures, v5: AAAI-26 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11231 2025-11-17 cs.CV cs.AI 83%

3D Gaussian and Diffusion-Based Gaze Redirection

Abiram Panchalingam, Indu Bodala, Stuart Middleton

机构 * School of Electronics and Computer Science, University of Southampton(电子与计算机科学学院,南安普顿大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08207 2025-11-14 cs.CV cs.LG 83%

DICE: Discrete Inversion Enabling Controllable Editing for Multinomial Diffusion and Masked Generative Models

Xiaoxiao He, Quan Dao, Ligong Han, Song Wen, Minhao Bai, Di Liu, Han Zhang, Martin Renqiang Min, Felix Juefei-Xu, Chaowei Tan, Bo Liu, Kang Li, Hongdong Li, Junzhou Huang, Faez Ahmed, Akash Srivastava, Dimitris Metaxas

机构 * Rutgers University(新泽西罗格斯大学) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室) Red Hat AI Innovation(红帽AI创新) Google DeepMind(谷歌DeepMind) NYU(纽约大学) Walmart Global Tech(沃尔玛全球技术) NEC Labs America(NEC美国实验室) Massachusetts Institute of Technology(麻省理工学院) ANU(澳大利亚国立大学) UT Arlington(德克萨斯大学阿灵顿分校)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Project webpage: https://hexiaoxiao-cs.github.io/DICE/. This paper was accepted to CVPR 2025 but later desk-rejected post camera-ready, due to a withdrawal from ICLR made 14 days before reviewer assignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08090 2025-11-12 cs.CV cs.AI 83%

StableMorph: High-Quality Face Morph Generation with Stable Diffusion

Wassim Kabbani, Kiran Raja, Raghavendra Ramachandra, Christoph Busch

机构 * Norwegian University of Science and Technology(挪威科学技术大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

Journal ref International Joint Conference on Biometrics 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07806 2025-11-12 cs.CV 83%

PC-Diffusion: Aligning Diffusion Models with Human Preferences via Preference Classifier

Shaomeng Wang, He Wang, Xiaolu Wei, Longquan Dai, Jinhui Tang

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(计算机科学与工程学院,南京理工大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments 10 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07499 2025-11-12 cs.CV cs.AI 83%

Toward the Frontiers of Reliable Diffusion Sampling via Adversarial Sinkhorn Attention Guidance

Kwanyoung Kim

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Accepted to AAAI 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24434 2025-11-11 cs.LG cs.CV 83%

Graph Flow Matching: Enhancing Image Generation with Neighbor-Aware Flow Fields

Md Shahriar Rahim Siddiqui, Moshe Eliasof, Eldad Haber

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

Comments The 40th Annual AAAI Conference on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18877 2025-11-11 cs.CV cs.CR cs.LG 83%

Mitigating Sexual Content Generation via Embedding Distortion in Text-conditioned Diffusion Models

Jaesin Ahn, Heechul Jung

机构 * Department of Artificial Intelligence(人工智能系) Kyungpook National University(全北国立大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments NeurIPS 2025 accepted. Official code: https://github.com/amoeba04/des

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09935 2025-11-11 eess.IV cs.CV physics.med-ph 83%

Physics-informed DeepCT: Sinogram Wavelet Decomposition Meets Masked Diffusion

Zekun Zhou, Tan Liu, Bing Yu, Yanru Gong, Liu Shi, Qiegen Liu

机构 * School of Mathematics and Computer Sciences, Nanchang University, Nanchang, China(南昌大学数学与计算机科学学院) School of Information Engineering, Nanchang University, Nanchang, China(南昌大学信息工程学院) Key Laboratory of Advanced Medical Imaging and Intelligent Computing of Guizhou Province, China(贵州省先进医学影像与智能计算重点实验室)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04117 2025-11-07 cs.CV 83%

Tortoise and Hare Guidance: Accelerating Diffusion Model Inference with Multirate Integration

Yunghee Lee, Byeonghyun Pak, Junwha Hong, Hoseong Kim

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

Comments 21 pages, 8 figures. NeurIPS 2025. Project page: https://yhlee-add.github.io/THG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27171 2025-11-06 cs.CV cs.AI 83%

H2-Cache: A Novel Hierarchical Dual-Stage Cache for High-Performance Acceleration of Generative Diffusion Models

Mingyu Sung, Il-Min Kim, Sangseok Yun, Jae-Mo Kang

机构 * Department of Artificial Intelligence, Kyungpook National University(人工智能系,庆尚国立大学) Department of Electrical and Computer Engineering, Queen’s University(电气与计算机工程系,皇后大学) Department of Information and Communications Engineering, Pukyong National University(信息与通信工程系,浦项国立大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02462 2025-11-05 cs.CV 83%

KAO: Kernel-Adaptive Optimization in Diffusion for Satellite Image

Teerapong Panboonyuen

机构 * Chulalongkorn University(朱拉隆功大学)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15724 2025-11-05 cs.CV 83%

A Practical Investigation of Spatially-Controlled Image Generation with Transformers

Guoxuan Xia, Harleen Hanspal, Petru-Daniel Tudosiu, Shifeng Zhang, Sarah Parisot

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

Comments TMLR https://openreview.net/forum?id=loT6xhgLYK

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01175 2025-11-05 cs.CV 83%

Diffusion Transformer meets Multi-level Wavelet Spectrum for Single Image Super-Resolution

Peng Du, Hui Li, Han Xu, Paul Barom Jeon, Dongwook Lee, Daehyun Ji, Ran Yang, Feng Zhu

机构 * Samsung R&D Institute China Xi’an (SRCX)(三星中国研发中心西安(SRCX)) Samsung Electronics Co., LTD., South Korea(三星电子有限公司,韩国)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments ICCV 2025 Oral Paper

详情

展开后加载摘要…

URL PDF HTML 收藏