arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86714 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70159 篇

2602.09449 2026-02-11 cs.CV 83%

Look-Ahead and Look-Back Flows: Training-Free Image Generation with Trajectory Smoothing

前瞻与回顾流:基于轨迹平滑的无训练图像生成

Yan Luo, Henry Huang, Todd Y. Zhou, Mengyu Wang

机构 * Harvard AI and Robotics Lab, Harvard University(哈佛人工智能与机器人实验室,哈佛大学)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 本文提出两种无训练轨迹平滑方法,通过优化潜在空间中的生成路径,在多个数据集上显著提升图像生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06250 2026-02-10 cs.LG cs.CV 83%

Test-Time Iterative Error Correction for Efficient Diffusion Models

测试时迭代误差校正用于高效扩散模型

Yunshan Zhong, Weiqi Yan, Yuxin Zhang

机构 * School of Computer Science and Technology, Hainan University(海南大学计算机科学与技术学院) MAC Lab, Department of Artificial Intelligence, School of Informatics, Xiamen University(厦门大学信息学院人工智能系)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文提出迭代误差校正方法,通过测试时迭代优化减少扩散模型生成误差,提升生成质量与效率的平衡。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17873 2026-02-09 cs.CV 83%

Preserving Spectral Structure and Statistics in Diffusion Models

在扩散模型中保持谱结构和统计信息

Baohua Yan, Jennifer Kava, Qingyuan Liu, Xuan Di

机构 * Department of Civil Engineering and Engineering Mechanics, Columbia University, NY, United States(土木工程与工程力学系,哥伦比亚大学,纽约,美国)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 PreSS通过在谱空间中保持信息性先验和结构信号,提升扩散模型的生成质量和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19254 2026-02-06 cs.CV 83%

Imperceptible Protection against Style Imitation from Diffusion Models

难以察觉的对抗风格模仿保护方法

Namhyuk Ahn, Wonhyuk Ahn, KiYoon Yoo, Daesik Kim, Seung-Hun Nam

机构 * Department of Electrical and Computer Engineering, Inha University(电子与计算机工程系,inha大学) NAVER WEBTOON AI KRAFTON

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文提出了一种在不降低保护效果的前提下提升图像保护不可察觉性的方法,通过感知图和难度感知保护机制优化保护强度。

Comments IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04406 2026-02-05 cs.CV 83%

LCUDiff: Latent Capacity Upgrade Diffusion for Faithful Human Body Restoration

LCUDiff: 基于潜在空间升级的扩散模型用于忠实的人体图像修复

Jue Gong, Zihan Zhou, Jingkai Wang, Shu Li, Libo Liu, Jianliang Lan, Yulun Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shenzhen Transsion Holdings Co., Ltd.(深圳Transsion控股有限公司)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 LCUDiff通过升级潜在空间和改进解码器路由,实现了更高质量的人体图像修复,提升了保真度和效率。

Comments 8 pages, 7 figures. The code and model will be at https://github.com/gobunu/LCUDiff

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21380 2026-02-05 cs.LG cs.CV 83%

Sparse-to-Sparse Training of Diffusion Models

扩散模型的稀疏到稀疏训练

Inês Cardoso Oliveira, Decebal Constantin Mocanu, Luis A. Leiva

机构 * University of Luxembourg(卢森堡大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 本文提出稀疏到稀疏训练范式,用于提升扩散模型在训练和推理效率上的性能,通过实验表明稀疏DMs在性能上优于密集模型,同时减少计算资源消耗。

Comments Accepted to TMLR

Journal ref Transactions on Machine Learning Research (TMLR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03510 2026-02-04 cs.CV 83%

Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers

语义路由:探索多层LLM特征加权用于扩散变换器

Bozhou Li, Yushuo Guan, Haolin Li, Bohan Zeng, Yiyan Ji, Yue Ding, Pengfei Wan, Kun Gai, Yuanxing Zhang, Wentao Zhang

机构 * Peking University(北京大学) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Nanjing University(南京大学) Fudan University(复旦大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 本文提出深度-wise语义路由方法,通过多层LLM特征加权提升文本到图像生成的准确性与稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03210 2026-02-04 cs.CV 83%

VIRAL: Visual In-Context Reasoning via Analogy in Diffusion Transformers

VIRAL: 通过类比在扩散变换器中实现视觉上下文推理

Zhiwen Li, Zhongjie Duan, Jinyan Ye, Cen Chen, Daoyuan Chen, Yaliang Li, Yingda Chen

机构 * School of Data Science and Engineering, East China Normal University, Shanghai, China(数据科学与工程学院,东华大学,上海,中国) Alibaba Group, Hangzhou, China(阿里巴巴集团,杭州,中国)

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV

AI总结 VIRAL通过类比和扩散变换器实现视觉上下文推理,解决视觉任务异质性问题,提升开放域编辑性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01077 2026-02-04 cs.CV 83%

PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers

PISA:分块稀疏注意力使扩散变换器更高效

Haopeng Li, Shitong Shao, Wenliang Zhong, Zikai Zhou, Lichen Bai, Hui Xiong, Zeke Xie

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 PISA通过分块稀疏注意力在保持质量的同时显著提升扩散变换器的效率。

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19154 2026-02-03 eess.IV cs.AI cs.CV 83%

RDDM: Practicing RAW Domain Diffusion Model for Real-world Image Restoration

RDDM: 在真实世界图像恢复中实践RAW域扩散模型

Yan Chen, Yi Wen, Wei Li, Junchao Liu, Yong Guo, Jie Hu, Xinghao Chen

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) Max Planck Institute for Informatics(马克斯·普朗克信息研究所)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 RDDM通过在RAW域直接恢复图像,克服了传统sRGB域扩散模型在高保真度与图像生成间的矛盾,提升了真实世界图像恢复的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00350 2026-02-03 cs.CV 83%

ReLAPSe: Reinforcement-Learning-trained Adversarial Prompt Search for Erased concepts in unlearned diffusion models

ReLAPSe: 通过强化学习训练的对抗性提示搜索消除未学习扩散模型中的擦除概念

Ignacy Kolton, Kacper Marzol, Paweł Batorski, Marcin Mazur, Paul Swoboda, Przemysław Spurek

机构 * Jagiellonian University(雅盖隆大学) IDEAS Research Institute(IDEAS研究机构)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 ReLAPSe通过强化学习训练对抗性提示搜索,有效恢复扩散模型中被移除的概念,提供高效的去学习测试工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22135 2026-01-30 cs.CV 83%

PI-Light: Physics-Inspired Diffusion for Full-Image Relighting

PI-Light:基于物理的扩散用于全图像重照明

Zhexin Liang, Zhaoxi Chen, Yongwei Chen, Tianyi Wei, Tengfei Wang, Xingang Pan

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV

AI总结 PI-Light通过结合物理引导的神经渲染模块和基于物理的损失函数,提升全图像重照明的泛化能力与现实场景适应性。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21517 2026-01-30 cs.CV 83%

HERS: Hidden-Pattern Expert Learning for Risk-Specific Vehicle Damage Adaptation in Diffusion Models

HERS:隐藏模式专家学习用于扩散模型中特定风险车辆损伤适应

Teerapong Panboonyuen

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 HERS通过领域特定专家学习提升扩散模型中车辆损伤生成的保真度与可控性,提高保险领域应用的安全性与可靠性。

Comments 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20766 2026-01-30 cs.CV 83%

DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion

DyPE: 用于超高清扩散的动态位置外推

Noam Issachar, Guy Yariv, Sagie Benaim, Yossi Adi, Dani Lischinski, Raanan Fattal

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 DyPE通过动态调整位置编码实现超高清扩散图像生成,无需额外训练成本,显著提升高分辨率图像保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20260 2026-01-29 cs.CV 83%

Reversible Efficient Diffusion for Image Fusion

可逆高效扩散用于图像融合

Xingxin Xu, Bing Cao, DongDong Li, Qinghua Hu, Pengfei Zhu

机构 * School of New Media and Communication, Tianjin University(新媒体与传播学院,天津大学) School of Artificial Intelligence, Tianjin University(人工智能学院,天津大学) Low-Altitude Intelligence Laboratory, Xiong’an National Innovation Center(低空智能实验室,雄安国家创新中心) Xiong’an Guochuang Lantian Technology Co., Ltd.(雄安国创蓝天科技有限公司) National University of Defense Technology(国防科技大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文提出可逆高效扩散模型,通过显式监督提升图像融合的生成能力,解决传统扩散模型在融合任务中的细节损失问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16324 2026-01-29 cs.CV 83%

From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation

从预测到完美:引入精修到自回归图像生成

Cheng Cheng, Lin Song, Di An, Yicheng Xiao, Xuchong Zhang, Hongbin Sun, Ying Shan

机构 * Xi’an Jiaotong University(西安交通大学) Johns Hopkins University(约翰霍普金斯大学) Tsinghua University(清华大学) ARC Lab, Tencent PCG(腾讯PCG ARC实验室)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 TensorAR通过引入张量预测机制,改进自回归图像生成的质量和性能。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17124 2026-01-28 cs.CV 83%

iFSQ: Improving FSQ for Image Generation with 1 Line of Code

iFSQ:通过一行代码改进图像生成的FSQ

Bin Lin, Zongjian Li, Yuwei Niu, Kaixiong Gong, Yunyang Ge, Yunlong Lin, Mingzhe Zheng, JianWei Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Li Yuan

机构 * Peking University(北京大学) Tencent Hunyuan(腾讯文言)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 iFSQ通过一行代码改进图像生成的FSQ,解决了重建保真度与信息效率的平衡问题,并揭示了离散与连续表示的最佳平衡点及AR与扩散模型的性能差异。

Comments Technical Report; Fixed eq.7 & 8 and corresponding content

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11194 2026-01-28 cs.LG cs.CV 83%

Beyond Memorization: Selective Learning for Copyright-Safe Diffusion Model Training

超越记忆:面向版权安全的扩散模型训练中的选择性学习

Divya Kothandaraman, Jaclyn Pytlarz

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 本文提出了一种基于梯度投影的选择性学习方法,用于在扩散模型训练中有效控制记忆化,以保障版权安全和隐私保护。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10495 2026-01-28 cs.CR cs.AI cs.CV cs.LG 83%

SWA-LDM: Toward Stealthy Watermarks for Latent Diffusion Models

SWA-LDM:迈向潜在扩散模型的隐蔽水印

Zhonghao Yang, Linye Lyu, Xuanhang Chang, Daojing He, YU LI

机构 * Software Engineering Institute, East China Normal University(华东师范大学软件工程学院) School of Computer Science and Technology, Harbin Institute of Technology (Shen Zhen)(哈尔滨工业大学(深圳)计算机科学与技术学院) College of Integrated Circuits, Zhejiang University(浙江大学集成电路学院)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 SWA-LDM通过动态随机化潜在噪声中的水印,提升潜在扩散模型的隐蔽性,实现更安全的水印部署。

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14738 2026-01-22 cs.CV 83%

Safeguarding Facial Identity against Diffusion-based Face Swapping via Cascading Pathway Disruption

通过级联路径破坏保障面部身份免受基于扩散的面部交换攻击

Liqin Wang, Qianyue Hu, Wei Lu, Xiangyang Luo

机构 * MoE Key Laboratory of Information Technology, Sun Yat-sen University, Guangzhou, China(信息科技教育部重点实验室,中山大学,广州,中国) State Key Laboratory of Mathematical Engineering and Advanced Computing, Zhengzhou, China(数学工程与先进计算国家重点实验室,郑州,中国)

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV

AI总结 VoidFace通过级联路径破坏技术,有效防御基于扩散模型的面部交换攻击,提升身份安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13416 2026-01-21 cs.CV 83%

Diffusion Representations for Fine-Grained Image Classification: A Marine Plankton Case Study

扩散表示用于细粒度图像分类:一个海洋浮游生物案例研究

A. Nieto Juscafresa, Á. Mazcuñán Herreros, J. Sullivan

机构 * KTH Royal Institute of Technology(皇家理工学院)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 本文提出利用冻结的扩散模型作为特征编码器,通过多层特征和时间步探测提升细粒度图像分类性能,在浮游生物监测中验证了其有效性。

Comments 21 pages, 6 figures, CVPR format

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11689 2026-01-21 eess.IV cs.CV 83%

Bridging Modalities: Joint Synthesis and Registration Framework for Aligning Diffusion MRI with T1-Weighted Images

弥合模态:联合合成与配准框架用于对齐扩散磁共振成像与T1加权图像

Xiaofan Wang, Junyi Wang, Yuqian Chen, Lauren J. O' Donnell, Fan Zhang

机构 * University of Electronic Science and Technology of China(电子科技大学) Brigham and Women’s Hospital(布里奇沃特医院) Harvard Medical School(哈佛医学院)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 本文提出一种基于生成配准网络的无监督框架,通过图像合成和变形场学习,提升扩散MRI与T1加权图像的配准精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11085 2026-01-19 eess.IV cs.CV physics.med-ph 83%

Generation of Chest CT pulmonary Nodule Images by Latent Diffusion Models using the LIDC-IDRI Dataset

利用LIDC-IDRI数据集通过潜在扩散模型生成胸部CT肺结节图像

Kaito Urata, Maiko Nagao, Atsushi Teramoto, Kazuyoshi Imaizumi, Masashi Kondo, Hiroshi Fujita

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 本文提出利用潜在扩散模型生成胸部CT肺结节图像,通过LIDC-IDRI数据集验证,生成图像质量与真实图像相当。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09213 2026-01-15 cs.CV cs.AI 83%

SpikeVAEDiff: Neural Spike-based Natural Visual Scene Reconstruction via VD-VAE and Versatile Diffusion

SpikeVAEDiff: 通过VD-VAE和多功能扩散实现神经尖峰基于的自然视觉场景重建

Jialu Li, Taiyan Zhou

机构 * HKUST Clear Water Bay(香港科技大学清水湾分校)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 SpikeVAEDiff通过VD-VAE和多功能扩散模型,利用神经尖峰数据实现高分辨率自然视觉场景重建,提升时间与空间分辨率,验证了特定脑区对重建质量的关键作用。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07530 2026-01-15 cs.CV 83%

Universal Few-Shot Spatial Control for Diffusion Models

通用少样本空间控制用于扩散模型

Kiet T. Nguyen, Chanhyuk Lee, Donggyun Kim, Dong Hoon Lee, Seunghoon Hong

机构 * KAIST(韩国科学技术院)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 UFC通过通用少样本控制适配器实现对扩散模型的通用空间控制,能够在少量示例下实现与全监督基线相当的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08095 2026-01-14 cs.CV 83%

From Prompts to Deployment: Auto-Curated Domain-Specific Dataset Generation via Diffusion Models

从提示到部署:通过扩散模型实现自动定制的领域特定数据集生成

Dongsik Yoon, Jongeun Kim

机构 * HDC LABS(HDC实验室)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

AI总结 本文提出通过扩散模型生成领域特定合成数据集的方法,解决预训练模型与现实部署间的分布偏移问题,通过三阶段框架高效构建高质量可部署数据集。

Comments To appear in the Workshop on Synthetic & Adversarial ForEnsics (SAFE), WACV 2026 (oral presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23130 2026-01-14 cs.CV cs.AI 83%

PathoSyn: Imaging-Pathology MRI Synthesis via Disentangled Deviation Diffusion

PathoSyn: 通过解耦偏差扩散实现影像-病理MRI合成

Jian Wang, Sixing Rong, Jiarui Xing, Yuling Xu, Weide Liu

机构 * Department of Radiology, Boston Children’s Hospital, Harvard Medical School(放射科,波士顿儿童医院,哈佛医学院) College of Science, Northeastern University(科学学院,东北大学) School of Medicine, Yale University(医学院,耶鲁大学) Department of Cardiac Surgery, The Second Affiliated Hospital of Jiangxi Medical College, Nanchang University(心脏外科科,江西医学院第二附属医院,南昌大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 PathoSyn通过解耦偏差扩散模型生成高保真MRI合成数据,提升低数据环境下诊断算法的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09644 2026-01-14 cs.CV cs.AI 83%

DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning

DGAE:基于扩散的自编码器用于高效潜在表示学习

Dongxu Liu, Jiahui Zhu, Yuang Peng, Haomiao Tang, Yuwei Chen, Chunrui Han, Zheng Ge, Daxin Jiang, Mingxue Liao

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Tsinghua University(清华大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 DGAE通过结合扩散模型提升解码器表达能力,实现高效潜在表示学习,在高压缩率下提升性能并减少潜在空间维度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07273 2026-01-13 cs.CV 83%

GenDet: Painting Colored Bounding Boxes on Images via Diffusion Model for Object Detection

GenDet:通过扩散模型在图像上绘制彩色边界框以实现目标检测

Chen Min, Chengyang Li, Fanjie Kong, Qi Zhu, Dawei Zhao, Liang Xiao

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 GenDet通过扩散模型将目标检测转化为图像生成任务,实现精确的边界框生成与语义标注,提升检测精度与灵活性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08635 2026-01-12 cs.CV 83%

Latent Diffusion Autoencoders: Toward Efficient and Meaningful Unsupervised Representation Learning in Medical Imaging

潜在扩散自编码器:迈向医学影像中高效且有意义的无监督表示学习

Gabriele Lozupone, Alessandro Bria, Francesco Fontanella, Frederick J. A. Meijer, Claudio De Stefano, Henkjan Huisman

机构 * Department of Electrical and Information Engineering (DIEI), University of Cassino and Southern Lazio(电气与信息工程系(DIEI),卡斯诺和南部拉齐亚大学) Diagnostic Image Analysis Group, Radboud University Medical Center(影像诊断分析组,拉德堡德大学医学中心) Department of Medical Imaging, Radboud University Medical Center(医学成像系,拉德堡德大学医学中心)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 LDAE通过在潜在空间中应用扩散过程,实现了高效且有意义的医学影像无监督学习,展示了在AD诊断和年龄预测中的高准确性及重建质量。

Comments 15 pages, 9 figures, 7 tables

Journal ref Medical Image Analysis (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏