arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86623 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70082 篇

2401.08741 2024-01-18 cs.CV cs.AI cs.LG 84%

Fixed Point Diffusion Models

Xingjian Bai, Luke Melas-Kyriazi

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Project page: https://lukemelas.github.io/fixed-point-diffusion-models

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.08332 2024-01-15 cs.CV 84%

Versatile Diffusion: Text, Images and Variations All in One Diffusion Model

Xingqian Xu, Zhangyang Wang, Eric Zhang, Kai Wang, Humphrey Shi

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments ICCV 2023; Github link: https://github.com/SHI-Labs/Versatile-Diffusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03221 2024-01-09 cs.CV cs.AI 84%

MirrorDiffusion: Stabilizing Diffusion Process in Zero-shot Image Translation by Prompts Redescription and Beyond

Yupei Lin, Xiaoyu Xian, Yukai Shi, Liang Lin

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments A prompt re-description strategy is proposed for stabilizing the diffusion model in image-to-image translation. Code and dataset page: https://mirrordiffusion.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00858 2023-12-11 cs.CV cs.AI 84%

DeepCache: Accelerating Diffusion Models for Free

Xinyin Ma, Gongfan Fang, Xinchao Wang

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

Comments Work in progress. Project Page: https://horseee.github.io/Diffusion_DeepCache/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04410 2023-12-08 cs.CV 84%

Smooth Diffusion: Crafting Smooth Latent Spaces in Diffusion Models

Jiayi Guo, Xingqian Xu, Yifan Pu, Zanlin Ni, Chaofei Wang, Manushree Vasu, Shiji Song, Gao Huang, Humphrey Shi

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments GitHub: https://github.com/SHI-Labs/Smooth-Diffusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00974 2023-12-01 cs.CV 84%

Discovering Failure Modes of Text-guided Diffusion Models via Adversarial Search

Qihao Liu, Adam Kortylewski, Yutong Bai, Song Bai, Alan Yuille

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Project page: https://sage-diffusion.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14303 2023-11-14 cs.CV 84%

Dataset Diffusion: Diffusion-based Synthetic Dataset Generation for Pixel-Level Semantic Segmentation

Quang Nguyen, Truong Vu, Anh Tran, Khoi Nguyen

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2023. Our project page: https://dataset-diffusion.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05720 2023-11-07 cs.CV cs.AI cs.LG 84%

Beyond Surface Statistics: Scene Representations in a Latent Diffusion Model

Yida Chen, Fernanda Viégas, Martin Wattenberg

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

Comments A short version of this paper is accepted in the NeurIPS 2023 Workshop on Diffusion Models: https://nips.cc/virtual/2023/74894

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28489 2026-07-31 cs.IT math.IT 新提交 83%

Forecasting Land Art Under Climate Scenarios

气候情景下的大地艺术预测

Alev Cinbarci, Sean Kalaycioglu

专题命中 扩散模型 :diffusion(summary_cn,abstract);image synthesis(abstract)

AI总结 本文基于《螺旋防波堤》的遥感数据,构建两阶段预测流程,结合IPCC气候情景与Stable Diffusion XL模型,预测该大地艺术的图像复杂性及暴露状态,并探讨文化遗产伦理问题。

Comments 15 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14075 2026-05-27 cs.LG cs.AI 83%

CFG-OEC: Classifier Free Guidance with Orthogonal Error Correction

CFG-OEC: 带正交误差校正的无分类器引导

Nakgyu Yang, Yechan Lee, SooJean Han

机构 * School of Electrical Engineering, Korea Advanced Institute of Science(韩国科学技术院电子工程学院)

专题命中 扩散模型 :diffusion(summary_cn,abstract);image generation(abstract)

AI总结 针对扩散模型中无分类器引导的采样规则与训练目标不匹配导致的误差,提出正交误差校正方法(CFG-OEC)通过减少条件与无条件预测误差的交互项来提升采样质量,并在Stable Diffusion上验证了FID和CLIP分数的改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20515 2026-08-24 cs.CV 新提交 83%

DiffVC-ONE: Diffusion-based Generative Video Compression with One-Step Video Diffusion Transformer

DiffVC-ONE:基于扩散模型的生成式视频压缩,采用单步视频扩散Transformer

Wenzhuo Ma, Zhenzhong Chen

机构 * school of Remote Sensing and Information Engineering, Wuhan University(武汉大学遥感信息工程学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DiffVC-ONE是基于扩散模型的生成式视频压缩框架,采用单步视频扩散Transformer,通过三类组件实现低推理成本下的高感知质量与时间一致性,在多基准上达最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19504 2026-08-21 cs.CV 新提交 83%

A Plug-in Interpretation of Conditioning in Score-Based Diffusion Models

基于分数的扩散模型中条件作用的插件式解释

Libo Chen, Souvik Ghosh, Teo Deveney, Chris Budd, Vinay P. Namboodiri

机构 * International Institute of Information Technology Hyderabad(国际信息技术学院海得拉巴分校) University of Bath(巴斯大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 该研究提出基于多速度联合扩散的插件式条件机制,推导相关SDE与ODE并引入对数福克-普朗克残差正则化,在条件图像生成任务中验证了方法的有效性。

Comments Accepted at BMVC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17747 2026-08-19 cs.CV 新提交 83%

TINA+: Probing Residual Visual Knowledge in Unlearned Diffusion Models via Diffusion-Consistent Text-Free Inversion

TINA+: 通过扩散一致无文本反演探测未学习的扩散模型中的残留视觉知识

Qianlong Xiang, Miao Zhang, Kun Wang, Haoyu Zhang, Junhui Hou, Liqiang Nie

机构 * School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)计算机科学与技术学院) City University of Hong Kong(香港城市大学) Shenzhen Loop Area Institute(深圳河套学院) National University of Singapore(新加坡国立大学) Pengcheng Laboratory(鹏城实验室)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 本研究提出TINA+,一种扩散一致无文本反演攻击,可探测未学习的扩散模型中残留的视觉知识,实验表明现有概念擦除方法多切断文本-图像链接而非消除视觉知识。

Comments The project page is https://qianlong0502.github.io/TINA-Plus-Homepage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16793 2026-08-18 cs.CV 新提交 83%

PixRestore: Unified Image Restoration via Pixel Diffusion Transformer

PixRestore:基于像素扩散Transformer的统一图像修复模型

Lingchen Sun, Rongyuan Wu, Xiangtao Kong, Jixin Zhao, Qiaosi Yi, Yujing Sun, Shuaizheng Liu, Zhengqiang Zhang, Lei Zhang

机构 * The Hong Kong Polytechnic University(香港理工大学) OPPO Research Institute(OPPO研究院)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 本文提出无VAE的像素空间DiT模型PixRestore,通过流匹配与DINO特征可靠性预测实现UIR,仅50M参数且单步推理,在效率与修复性能上优于同类模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15452 2026-08-18 cs.CV 新提交 83%

Spatially-Grounded Flow Matching: Structured Source Distributions for Image Generation

空间接地流匹配:用于图像生成的结构化源分布

Arman Zarei, Mahdi M. Kalayeh

机构 * University of Maryland(马里兰大学) Netflix(网飞公司)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 针对流匹配模型源分布缺乏空间结构的问题,提出StructFlow将空间局部性编码到源分布,可提升图像生成质量与局部可控重合成性能,且能集成到大型预训练模型中。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03929 2026-08-17 cs.LG cs.CV 版本更新 83%

Latent Reward Registers for Diffusion Preference Alignment

用于扩散偏好对齐的潜在奖励寄存器

Yuanshen Guan, Zipeng Feng, Chengru Song, Zhiwei Xiong, Peiqin Sun

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对扩散模型偏好对齐的时间信用分配挑战,提出Latent Reward Registers机制,结合RG-OPD和RGS策略,在高噪声下实现最优性能且大幅降低计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06526 2026-08-17 cs.CV 版本更新 83%

Concept Unlearning by Modeling Key Steps of Diffusion Process

通过建模扩散过程关键步骤实现的概念遗忘

Chaoshuo Zhang, Chenhao Lin, Zhengyu Zhao, Le Yang, Qian Wang, Chao Shen

机构 * School of Cyber Science and Engineering, Xi’an Jiaotong University(网络安全科学与工程学院,西安交通大学) Wuhan University(武汉大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 针对文本到图像扩散模型概念遗忘时的灾难性遗忘问题,提出KSCU方法,通过动态隔离优化至概念特定活跃区域,实现了更优的概念擦除与效用保留权衡,在多类任务中性能领先。

Comments Accepted by T-IFS

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13455 2026-08-14 cs.CV 新提交 83%

Evaluation of Clinically Steerable Retinal Image Generation from Foundation Model Latent Spaces

基于基础模型潜在空间的临床可操控视网膜图像生成评估

Zuzanna A. Wakefield-Skórniewska, Bartłomiej W. Papież

机构 * Nuffield Department of Population Health, University of Oxford(牛津大学纳菲尔德人口健康系) University of Oxford(牛津大学)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 该研究评估了四种视网膜基础模型的可控图像生成能力,发现其在自身框架内生成的视网膜图像优于传统潜在扩散,但与真实图像存在表征差距,需进一步对齐。

Comments MICCAI 2026 Workshop SASHIMI Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13043 2026-08-14 cs.AI cs.CV cs.LG 新提交 83%

From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion

从局部失配到全局影响:优化缓存重用策略以实现高效扩散模型

Xichen Ye, Yifan Wu, Zhikang Xie, Xiangyu Yue, Cheng Jin, Weizhong Zhang

机构 * Faculty of Engineering, The Chinese University of Hong Kong(香港中文大学工程学院) School of Data Science, Fudan University(复旦大学数据科学学院)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 针对扩散模型缓存策略与生成质量失配问题,提出GCache双层优化框架,在Wan2.1模型上实现2.17倍加速且LPIPS降至0.0316,性能优于现有缓存策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08063 2026-08-14 cs.CV 版本更新 83%

Interpretable ODE-style Generative Diffusion Model via Force Field Construction

基于力场构造的可解释ODE式生成扩散模型

Weiyang Jin, Yongpei Zhu, Yuxi Peng

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文从数学视角识别适配的物理模型,归纳出统一的ODE式生成扩散模型方法,在CIFAR-10实验中验证了该方法在生成速度、Inception分数和FID分数上的优异性能,提升了扩散模型的可解释性。

Comments This is empirical reproducing and experimental tries of Jianlin Su's blog use ChatGPT and Newbeing: https://kexue.fm/archives/9370 https://kexue.fm/archives/9379 https://kexue.fm/archives/9467 We make the mistakes to cite them and this paper need to be withdrawn to clearify this. We thank Jianlin reached out us and give us chance to claim

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11807 2026-08-13 cs.CV 新提交 83%

CoDiR: Confidence-Guided Diffusion Refinement for Semi-Supervised Histopathology Segmentation

CoDiR:用于半监督组织病理学分割的置信度引导扩散精调

Hoai Nhan Pham, Dang-Nguyen Bui, Le-Van Thai, Thanh-Hiep Vo, Lan Anh Dinh Thi, Tien Dat Nguyen, Duy-Dong Nguyen, Ngoc Lam Quang Bui, Tam Tran, Zhi Huang

机构 * AI VIETNAM Lab(AI VIETNAM实验室) Portland Community College(波特兰社区学院) University of Science, VNU-HCM(胡志明市国家大学科学大学) Hanoi University of Science and Technology(河内科技大学) Jeonbuk National University(全北国立大学) Washington University School of Medicine(华盛顿大学医学院) Perelman School of Medicine, University of Pennsylvania(宾夕法尼亚大学佩雷尔曼医学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 针对半监督组织病理学分割的标注稀缺与伪标签不可靠问题,提出CoDiR框架,结合Mean Teacher与扩散模型精调伪标签,在GlaS、CRAG数据集上取得优异mDice指标,精调模块贡献显著

Comments Accepted to the MICCAI COMPAYL Workshop 2026 (11 pages, 2 figures, 6 tables)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17585 2026-08-13 cs.CV 版本更新 83%

Pixel-Space Diffusion Transformers

像素空间扩散变换器

Renye Yan, Jikang Cheng, You Wu, Ling Liang, Wei Peng, Athanasios V. Vasilakos, Qingyu Zhao, Yu Zhang, Yimao Cai, Kilian M. Pohl, Guoying Zhao

机构 * Peking University(北京大学) Nanjing University(南京大学) Stanford University(斯坦福大学) Cornell University(康奈尔大学) University of Oulu(奥卢大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 本文探讨像素空间扩散变换器,针对潜在扩散模型的局限,研究直接对原始像素建模的像素空间扩散方法,介绍其在高维建模中的挑战与多模态建模优势,从多方面回顾pDiTs,总结方法、识别挑战并展望未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10995 2026-08-12 cs.CV 新提交 83%

HNDiff: Haze-Noise Diffusion for Image Dehazing

HNDiff:用于图像去雾的雾霾-噪声扩散模型

Jin-Ting He, Fu-Jen Tsai, Yan-Tsung Peng, Min-Hung Chen, Chia-Wen Lin, Yen-Yu Lin

机构 * National Yang Ming Chiao Tung University(国立阳明交通大学) National Tsing Hua University(国立清华大学) National Chengchi University(国立政治大学) NVIDIA(英伟达)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出嵌入大气散射模型的HNDiff扩散框架,通过雾霾感知噪声调度器关联正向退化物理机制,反向同步去雾去噪,改进主流去雾骨干网络并在基准数据集达最优结果。

Comments Accepted to ECCV 2026. Project Page: https://jin-ting-he.github.io/HNDiff

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05746 2026-08-12 cs.CV 版本更新 83%

HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models

HQ-DM:基于单Hadamard变换的量化感知训练用于低比特扩散模型

Shizhuo Mao, Hongtao Zou, Qihu Xie, Song Chen, Yi Kang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 HQ-DM通过单Hadamard变换提升低比特扩散模型的量化性能,显著提高Inception Score

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02743 2026-08-11 cs.CV 版本更新 83%

MultiShadow: Multi-Object Shadow Generation for Image Compositing via Diffusion Model

MultiShadow: 基于扩散模型的多物体阴影生成用于图像合成

Waqas Ahmed, Dean Diepeveen, Ferdous Sohel

机构 * School of Information Technology, Murdoch University(信息科技学院,穆尔彻大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 本文提出基于扩散模型的多物体阴影生成方法,通过多模态特征融合和注意力对齐损失,实现多物体阴影的物理合理合成。

Comments This work is currently under consideration for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12763 2026-08-11 cs.CV 版本更新 83%

AMD:Anatomical Motion Diffusion with Interpretable Motion Decomposition and Fusion

AMD:结合可解释动作分解与融合的解剖学动作扩散模型

Beibei Jing, Youjia Zhang, Zikai Song, Junqing Yu, Wei Yang

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出AMD模型,利用LLM将文本解析为解剖学脚本,结合双分支融合方案平衡文本与脚本影响,在CLCD1、CLCD2等复杂动作数据集上性能优于现有SOTA模型。

Comments Corrected missing spaces in the abstract; no changes to the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07463 2026-08-10 cs.CV cs.LG 新提交 83%

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

MirrorWorld:利用视频扩散模型生成镜像反射

Youjun Zhao, Alex Warren, Gary K. L. Tam, Rynson W. H. Lau

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

AI总结 MirrorWorld是一种反射感知视频修复框架,通过语义关系蒸馏和几何变换对齐建模场景与镜像的关系,在视频镜像反射重建任务上优于现有方法。

Comments Project Page: https://youjunzhao.github.io/MirrorWorld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06794 2026-08-10 cs.CV 新提交 83%

PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model

PAST:用于高效扩散模型的提示自适应采样终止

Renye Yan, Jikang Cheng, You Wu, Wei Peng, Zongwei Wang, Ling Liang, Yimao Cai

机构 * Peking University(北京大学) Nanjing University(南京大学) Stanford University(斯坦福大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 PAST是一种提示自适应采样终止方法,通过双自适应协调机制提升扩散模型RL微调的计算效率与偏好优化质量,效率最高提升66.7%,质量最高提升29.5%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05878 2026-08-07 cs.CV 新提交 83%

MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers

MAVISEG:用于扩散Transformer中零样本开放词汇分割的流形传播与视觉原型

Rajatsubhra Chakraborty, Xujun Che, Ritabrata Chakraborty, Xi Niu, Depeng Xu

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 MAVISEG是一种无需训练的优化层,通过恢复扩散Transformer的结构化信号,在六个基准测试的无需训练方法中取得了最强整体结果,其mIoU在各基准中均最优。

Comments 20 pages, 14 figures, 9 tables. Preprint under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04394 2026-08-06 cs.CV 新提交 83%

Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection

通过重新审视基于扩散模型的跨域小样本目标检测数据生成实现无额外开销的数据增强

Zijian Zhuang, Yixiong Zou, Yuhua Li, Ruixuan Li

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

AI总结 针对跨域小样本目标检测的域差距与数据稀缺挑战,提出带定制噪声的选择性修复(SITN)方法,结合定制噪声生成与选择性修复模块合成有效数据,在6个CDFSOD和4个CDFSS数据集上达到新的最先进性能

详情

展开后加载摘要…

URL PDF HTML 收藏