arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2026-02-03 至 2026-02-03 共收录 115 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 115 篇

2602.00413 2026-02-03 stat.ML cs.LG 91%

Alignment of Diffusion Model and Flow Matching for Text-to-Image Generation

扩散模型与流匹配的对齐方法用于文本到图像生成

Yidong Ouyang, Liyan Xie, Hongyuan Zha, Guang Cheng

机构 * Department of Statistics, University of California, Los Angeles(加州大学洛杉矶分校统计学系) Department of Industrial and Systems Engineering, University of Minneasota(明尼苏达大学工业与系统工程系) School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院)

专题命中 扩散模型 :image generation(title,abstract);text-to-image(title,abstract);diffusion(title,abstract)

AI总结 本文提出了一种无需微调的对齐框架,通过分数指导和速度指导分别优化扩散模型和流匹配模型,显著降低计算成本并提升生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02114 2026-02-03 cs.CV cs.LG 88%

Enhancing Diffusion-Based Quantitatively Controllable Image Generation via Matrix-Form EDM and Adaptive Vicinal Training

通过矩阵形式EDM和自适应邻近训练增强基于扩散的可量化控制图像生成

Xin Ding, Yun Chen, Sen Zhang, Kao Zhang, Nenglun Chen, Peibei Cao, Yongwei Wang, Fei Wu

专题命中 扩散模型 :diffusion(title,abstract);image generation(title);text-to-image(abstract);分类 cs.CV

AI总结 本文提出改进的CCDM框架iCCDM,结合矩阵形式EDM和自适应邻近训练,提升图像生成质量和采样效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01609 2026-02-03 cs.CV 87%

Token Pruning for In-Context Generation in Diffusion Transformers

令牌剪枝用于扩散变换器中的上下文生成

Junqing Lin, Xingyu Zheng, Pei Cheng, Bin Fu, Jingwei Sun, Guangzhong Sun

机构 * University of Science and Technology of China, Anhui, China(中国科学技术大学) Beihang University, Beijing, China(北京航空航天大学) Tencent PCG, China(腾讯PCG)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image synthesis(abstract)

AI总结 ToPi通过无训练的令牌剪枝框架提升扩散变换器的上下文生成效率,实现30%以上的推理加速同时保持生成质量。

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02051 2026-02-03 cs.AI 85%

SIDiffAgent: Self-Improving Diffusion Agent

SIDiffAgent:自改进扩散代理

Shivank Garg, Ayush Singh, Gaurav Kumar Nayak

机构 * Indian Institute of Technology, Roorkee(印度理工学院罗尔基分校)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);image synthesis(abstract)

AI总结 SIDiffAgent通过自改进的扩散代理框架,利用Qwen系列模型提升文本到图像生成的可靠性与可控性,实现更高质量的图像输出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19154 2026-02-03 eess.IV cs.AI cs.CV 83%

RDDM: Practicing RAW Domain Diffusion Model for Real-world Image Restoration

RDDM: 在真实世界图像恢复中实践RAW域扩散模型

Yan Chen, Yi Wen, Wei Li, Junchao Liu, Yong Guo, Jie Hu, Xinghao Chen

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) Max Planck Institute for Informatics(马克斯·普朗克信息研究所)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

AI总结 RDDM通过在RAW域直接恢复图像,克服了传统sRGB域扩散模型在高保真度与图像生成间的矛盾,提升了真实世界图像恢复的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00350 2026-02-03 cs.CV 83%

ReLAPSe: Reinforcement-Learning-trained Adversarial Prompt Search for Erased concepts in unlearned diffusion models

ReLAPSe: 通过强化学习训练的对抗性提示搜索消除未学习扩散模型中的擦除概念

Ignacy Kolton, Kacper Marzol, Paweł Batorski, Marcin Mazur, Paul Swoboda, Przemysław Spurek

机构 * Jagiellonian University(雅盖隆大学) IDEAS Research Institute(IDEAS研究机构)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

AI总结 ReLAPSe通过强化学习训练对抗性提示搜索,有效恢复扩散模型中被移除的概念,提供高效的去学习测试工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01362 2026-02-03 cs.CL 82%

Balancing Understanding and Generation in Discrete Diffusion Models

在离散扩散模型中实现理解和生成的平衡

Yue Liu, Yuzhong Zhao, Zheyong Xie, Qixiang Ye, Jianbin Jiao, Yao Hu, Shaosheng Cao, Yunfan Liu

机构 * Xiaohongshu Inc.(小红书公司)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

AI总结 XDLM通过平稳噪声核平衡理解和生成能力,实现MDLM和UDLM的理论统一并提升生成质量与理解能力。

Comments 32 pages, Code is available at https://github.com/MzeroMiko/XDLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02107 2026-02-03 cs.CV 79%

Teacher-Guided Student Self-Knowledge Distillation Using Diffusion Model

教师引导的学生自知识蒸馏使用扩散模型

Yu Wang, Chuanguang Yang, Zhulin An, Weilun Feng, Jiarui Zhao, Chengqing Yu, Libo Huang, Boyu Diao, Yongjun Xu

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 教师引导的学生自知识蒸馏使用扩散模型,通过轻量级扩散模型和局部敏感哈希方法提升学生模型对教师知识的学习效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02092 2026-02-03 cs.CV 79%

FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space

FSVideo: 一种在高度压缩潜在空间中的快速速度视频扩散模型

FSVideo Team, Qingyu Chen, Zhiyuan Fang, Haibin Huang, Xinwei Huang, Tong Jin, Minxuan Lin, Bo Liu, Celong Liu, Chongyang Ma, Xing Mei, Xiaohui Shen, Yaojie Shen, Fuwen Tan, Angtian Wang, Xiao Yang, Yiding Yang, Jiamin Yuan, Lingxi Zhang, Yuxin Zhang

机构 * FSVideo Team Intelligent Creation(FSVideo团队智能创作)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 FSVideo通过高度压缩的潜在空间和改进的扩散Transformer架构,在速度和质量上实现了高效视频生成。

Comments Project Page: https://kingofprank.github.io/fsvideo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01949 2026-02-03 cs.LG cs.CV 79%

Boundary-Constrained Diffusion Models for Floorplan Generation: Balancing Realism and Diversity

边界约束扩散模型用于布局生成:在真实感与多样性之间平衡

Leonardo Stoppani, Davide Bacciu, Shahab Mokarizadeh

机构 * University of Pisa(比萨大学) H&M Group(H&M集团)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出边界约束扩散模型,通过引入BCA模块提升边界一致性,并引入DS指标平衡布局生成中的真实感与多样性。

Comments Accepted at ESANN 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01570 2026-02-03 cs.CV 79%

One-Step Diffusion for Perceptual Image Compression

单步扩散用于感知图像压缩

Yiwen Jia, Hao Wei, Yanhui Zhou, Chenyang Ge

机构 * Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所) School of Information and telecommunication(信息与电信学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出一种单步扩散图像压缩方法,通过紧凑特征表示和判别器提升感知质量,实现46倍更快的推理速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18623 2026-02-03 cs.CV 79%

Adaptive Domain Shift in Diffusion Models for Cross-Modality Image Translation

扩散模型中自适应域偏移用于跨模态图像翻译

Zihao Wang, Yuzhou Chen, Shaogang Ren

机构 * Laplace Lab at University of Tennessee(田纳西大学Laplace实验室) University of California Riverside(加州大学河滨分校) University of Tennessee at Chattanooga(田纳西大学查塔努加分校)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出了一种自适应域偏移方法,通过在扩散模型生成过程中嵌入域偏移动态,提升跨模态图像翻译的结构保真度和语义一致性。

Comments Paper accepted as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04456 2026-02-03 cs.CV cs.AI 79%

GuidNoise: Single-Pair Guided Diffusion for Generalized Noise Synthesis

GuidNoise:单对引导扩散用于通用噪声合成

Changjin Kim, HyeokJun Lee, YoungJoon Yoo

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 GuidNoise通过单对引导扩散模型实现通用噪声合成,无需额外元数据,提升去噪性能,尤其适用于轻量模型和有限数据场景。

Comments AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13745 2026-02-03 cs.CV 79%

UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese Calligraphy

UniCalli: 一种统一的扩散框架,用于中文书法的列级生成与识别

Tianshuo Xu, Kai Wang, Zhifei Chen, Leyi Wu, Tianshui Wen, Fei Chao, Ying-Cong Chen

机构 * HKUST(GZ)(香港科技大学(广州)) China University of Geoscience Beijing(中国地质大学(北京)) Xiamen University(厦门大学) HKUST(香港科技大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 UniCalli提出一种统一的扩散框架,通过联合训练实现中文书法列级生成与识别,提升生成质量和识别性能,同时扩展至其他古代文字。

Comments Page: https://envision-research.github.io/UniCalli/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20295 2026-02-03 cs.CV 79%

FAST: Foreground-aware Diffusion with Accelerated Sampling Trajectory for Segmentation-oriented Anomaly Synthesis

FAST: 前景感知扩散与加速采样轨迹用于面向分割的异常合成

Xichen Xu, Yanshu Wang, Jinbao Wang, Xiaoning Lei, Guoyang Xie, Guannan Jiang, Zhichao Lu

机构 * Global Institute of Future Technology, Shanghai Jiao Tong University, Shanghai, China(上海交通大学未来技术全球研究院) School of Artificial Intelligence, Shenzhen University, Shenzhen, China(深圳大学人工智能学院) Department of Intelligent Manufacturing, CATL, Ningde, China(CATL智能制造部门) Department of Computer Science, City University of Hong Kong, Hong Kong, China(香港城市大学计算机科学系)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 FAST提出一种前景感知扩散框架,通过AIAS和FARM模块提升工业异常合成效率与质量,实现更可控的结构特定异常生成。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10701 2026-02-03 cs.CV cs.LG eess.IV 79%

Diffusion-based Layer-wise Semantic Reconstruction for Unsupervised Out-of-Distribution Detection

基于扩散的逐层语义重建用于无监督分布外检测

Ying Yang, De Cheng, Chaowei Fang, Yubiao Wang, Changzhe Jiao, Lechao Cheng, Nannan Wang

机构 * Xidian University(西电大学) Hefei University of Technology(合肥工业大学) Chongqing University of Posts and Telecommunications(重庆邮电大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出基于扩散的逐层语义重建方法,用于无监督分布外检测,通过特征重建误差区分ID和OOD样本,实现高准确性和效率。

Comments 26 pages, 23 figures, published to Neurlps2024

Journal ref Proceedings of the 38th Conference on Neural Information Processing Systems (NeurIPS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01303 2026-02-03 cs.CV 79%

ReDiStory: Region-Disentangled Diffusion for Consistent Visual Story Generation

ReDiStory: 区域解耦扩散用于一致的视觉故事生成

Ayushman Sarkar, Zhenyu Yu, Chu Chen, Wei Tang, Kangning Cui, Mohd Yamani Idna Idris

机构 * Birbhum Institute of Engineering and Technology(比尔布尔工程科技学院) Universiti Malaya(马来亚大学) City University of Hong Kong(香港城市大学) Wake Forest University(威克森林大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 ReDiStory通过推理时的提示嵌入重组,提升多帧视觉故事生成中主体身份的一致性,无需修改扩散参数或额外监督。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00839 2026-02-03 cs.CV 79%

TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation

TransNormal: 用于扩散基透明物体法线估计的密集视觉语义

Mingwei Li, Hehe Fan, Yi Yang

机构 * Zhejiang University(浙江大学) Zhongguancun Academy(中关村学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 TransNormal通过整合密集视觉语义和多任务学习,提升透明物体法线估计的精度与鲁棒性。

Comments Project Page: https://longxiang-ai.github.io/TransNormal

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00739 2026-02-03 cs.CV 79%

Diffusion-Driven Inter-Outer Surface Separation for Point Clouds with Open Boundaries

基于扩散的点云双层表面分离:用于开放边界

Zhengyan Qin, Liyuan Qiu

机构 * Hong Kong University of Science and Technology (HKUST)(香港理工大学) Hong Kong Applied Science and Technology Research Institute (ASTRI)(香港应用科技研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于扩散的算法,用于分离双层点云的内层和外层表面,特别针对具有开放边界的点云,通过提取真实内层来解决重叠表面和法线紊乱问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00583 2026-02-03 cs.CV cs.AI 79%

MAUGen: A Unified Diffusion Approach for Multi-Identity Facial Expression and AU Label Generation

MAUGen: 一种用于多身份面部表情和AU标签生成的统一扩散方法

Xiangdong Li, Ye Lou, Ao Gao, Wei Zhang, Siyang Song

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 MAUGen通过统一的扩散方法生成多身份面部表情和AU标签,提升面部动作单元识别系统的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00536 2026-02-03 cs.CV 79%

SADER: Structure-Aware Diffusion Framework with DEterministic Resampling for Multi-Temporal Remote Sensing Cloud Removal

SADER:一种结构感知扩散框架,用于多时相遥感云去除

Yifan Zhang, Qian Chen, Yi Liu, Wengen Li, Jihong Guan

机构 * College of Literature, Science, and the Arts, University of Michigan(文学、科学与艺术学院,密歇根大学) School of Computer Science and Technology, Tongji University(计算机科学与技术学院,同济大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 SADER通过结构感知扩散框架,结合时间融合和混合注意力机制,有效解决多时相遥感云去除问题,提升云去除效果和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18537 2026-02-03 cs.CV 79%

Zero-Shot Video Deraining with Video Diffusion Models

无监督视频去雨与视频扩散模型

Tuomas Varanka, Juan Luis Gonzalez, Hyeongwoo Kim, Pablo Garrido, Xu Yao

机构 * University of Oulu(奥卢大学) Flawless AI Imperial College London(伦敦帝国理工学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出了一种无需合成数据和模型微调的无监督视频去雨方法,通过预训练文本到视频扩散模型,利用注意力切换机制提升动态场景中的去雨效果。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25178 2026-02-03 cs.CV cs.AI cs.LG 79%

GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs

GHOST:诱导幻觉的多模态大语言模型图像生成

Aryan Yazdan Parast, Parsa Hosseini, Hesam Asadollahzadeh, Arshia Soltani Moakhar, Basim Azam, Soheil Feizi, Naveed Akhtar

机构 * The University of Melbourne(墨尔本大学) University of Maryland(马里兰大学)

专题命中 扩散模型 :image generation(title);diffusion(abstract);分类 cs.CV

AI总结 GHOST通过优化隐蔽令牌生成诱导幻觉的图像,评估多模态大语言模型的可靠性,并发现高幻觉成功率及可转移漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01330 2026-02-03 cs.CV 79%

Prior-Guided Residual Diffusion: Calibrated and Efficient Medical Image Segmentation

先验引导残差扩散:校准且高效的医学图像分割

Fuyou Mao, Beining Wu, Yanfeng Jiang, Han Xue, Yan Tang, Hao Zhang

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 PGRD通过先验引导残差扩散方法,在医学图像分割中实现高精度与高效校准。

Comments Withdrawn by the authors to conduct further methodological refinement and address concerns regarding the originality of the current implementation

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15078 2026-02-03 eess.IV cs.CV physics.med-ph 79%

PET Image Reconstruction Using Deep Diffusion Image Prior

基于深度扩散图像先验的PET图像重建

Fumio Hashimoto, Kuang Gong

机构 * J. Crayton Pruitt Family Department of Biomedical Engineering, University of Florida(朱·克雷顿·普瑞特家庭生物医学工程系,佛罗里达大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出基于深度扩散模型的PET图像重建方法,通过解剖先验引导和半二次分裂算法实现高效重建,适用于多种示踪剂和扫描仪类型。

Comments 11 pages, 12 figures

Journal ref IEEE Trans. Med. Imaging (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23993 2026-02-03 cs.CV cs.AI 79%

DenseFormer: Learning Dense Depth Map from Sparse Depth and Image via Conditional Diffusion Model

DenseFormer: 通过条件扩散模型学习稀疏深度和图像的密集深度图

Ming Yuan, Chuang Zhang, Lei He, Qing Xu, Jianqiang Wang

机构 * the School of Vehicle and Mobility(车辆与移动学院) Tsinghua University(清华大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DenseFormer通过条件扩散模型,结合特征提取和深度细化模块,实现从稀疏深度和图像生成密集深度图,优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02238 2026-02-03 cs.LG cs.AI 78%

Geometry- and Relation-Aware Diffusion for EEG Super-Resolution

基于几何和关系的扩散模型用于EEG超分辨率

Laura Yao, Gengwei Zhang, Moajjem Chowdhury, Yunmei Liu, Tianlong Chen

机构 * Electroencephalography (EEG) provides a noninvasive window into distributed cortical dynamics of brain and is widely used across affective computing(脑电图(EEG)为研究大脑分布式皮层动态提供了一种非侵入性窗口,并广泛应用于情感计算)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出TopoDiff模型,通过结合拓扑感知嵌入和动态通道关系图,提升EEG空间超分辨率的生成保真度和下游任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02159 2026-02-03 cs.CL 78%

Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing

Focus-dLLM: 通过置信度引导的上下文聚焦加速长上下文扩散语言模型推理

Lingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong, Jun Zhang, Ao Zhou, Jianlei Yang

机构 * Beihang University(北航) Hong Kong University of Science and Technology(香港理工大学) SenseTime Research(商汤科技研究院)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 Focus-dLLM通过置信度引导的上下文聚焦技术,实现了长上下文扩散语言模型推理的高效加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02113 2026-02-03 stat.ML cs.LG cs.NA math.NA 78%

Training-free score-based diffusion for parameter-dependent stochastic dynamical systems

无训练的基于分数的扩散模型用于参数依赖的随机动力系统

Minglei Yang, Sicheng He

机构 * Fusion Energy Division, Oak Ridge National Laboratory(奥克兰国家实验室融合能源部门) Department of Mechanical and Aerospace Engineering, University of Tennessee(田纳西大学机械与航空航天工程系)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出了一种无训练的条件扩散模型,用于高效学习参数依赖随机微分方程的随机流映射,通过联合核加权蒙特卡洛估计器实现连续参数域的插值,提升参数研究和不确定性量化效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01909 2026-02-03 physics.geo-ph cs.LG 78%

Propagating the prior from far to near offset: A self-supervised diffusion framework for progressively recovering near-offsets of towed-streamer data

从远到近传播先验:一种自监督扩散框架用于逐步恢复拖曳线阵数据的近偏移

Shijun Cheng, Tariq Alkhalifah

机构 * Division of Physical Science and Engineering(物理科学与工程系) King Abdullah University of Science and Technology(国王阿卜杜勒·阿齐兹大学科学与技术学院)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出一种自监督扩散框架,通过递归外推从远到近恢复拖曳线阵数据的近偏移轨迹,无需地面真实数据,提升地震数据处理效果。

详情

展开后加载摘要…

URL PDF HTML 收藏