arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 70159 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70159 篇

2509.20295 2026-02-03 cs.CV 79%

FAST: Foreground-aware Diffusion with Accelerated Sampling Trajectory for Segmentation-oriented Anomaly Synthesis

FAST: 前景感知扩散与加速采样轨迹用于面向分割的异常合成

Xichen Xu, Yanshu Wang, Jinbao Wang, Xiaoning Lei, Guoyang Xie, Guannan Jiang, Zhichao Lu

机构 * Global Institute of Future Technology, Shanghai Jiao Tong University, Shanghai, China(上海交通大学未来技术全球研究院) School of Artificial Intelligence, Shenzhen University, Shenzhen, China(深圳大学人工智能学院) Department of Intelligent Manufacturing, CATL, Ningde, China(CATL智能制造部门) Department of Computer Science, City University of Hong Kong, Hong Kong, China(香港城市大学计算机科学系)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 FAST提出一种前景感知扩散框架,通过AIAS和FARM模块提升工业异常合成效率与质量,实现更可控的结构特定异常生成。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10701 2026-02-03 cs.CV cs.LG eess.IV 79%

Diffusion-based Layer-wise Semantic Reconstruction for Unsupervised Out-of-Distribution Detection

基于扩散的逐层语义重建用于无监督分布外检测

Ying Yang, De Cheng, Chaowei Fang, Yubiao Wang, Changzhe Jiao, Lechao Cheng, Nannan Wang

机构 * Xidian University(西电大学) Hefei University of Technology(合肥工业大学) Chongqing University of Posts and Telecommunications(重庆邮电大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出基于扩散的逐层语义重建方法,用于无监督分布外检测,通过特征重建误差区分ID和OOD样本,实现高准确性和效率。

Comments 26 pages, 23 figures, published to Neurlps2024

Journal ref Proceedings of the 38th Conference on Neural Information Processing Systems (NeurIPS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01303 2026-02-03 cs.CV 79%

ReDiStory: Region-Disentangled Diffusion for Consistent Visual Story Generation

ReDiStory: 区域解耦扩散用于一致的视觉故事生成

Ayushman Sarkar, Zhenyu Yu, Chu Chen, Wei Tang, Kangning Cui, Mohd Yamani Idna Idris

机构 * Birbhum Institute of Engineering and Technology(比尔布尔工程科技学院) Universiti Malaya(马来亚大学) City University of Hong Kong(香港城市大学) Wake Forest University(威克森林大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 ReDiStory通过推理时的提示嵌入重组,提升多帧视觉故事生成中主体身份的一致性,无需修改扩散参数或额外监督。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00839 2026-02-03 cs.CV 79%

TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation

TransNormal: 用于扩散基透明物体法线估计的密集视觉语义

Mingwei Li, Hehe Fan, Yi Yang

机构 * Zhejiang University(浙江大学) Zhongguancun Academy(中关村学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 TransNormal通过整合密集视觉语义和多任务学习,提升透明物体法线估计的精度与鲁棒性。

Comments Project Page: https://longxiang-ai.github.io/TransNormal

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00739 2026-02-03 cs.CV 79%

Diffusion-Driven Inter-Outer Surface Separation for Point Clouds with Open Boundaries

基于扩散的点云双层表面分离:用于开放边界

Zhengyan Qin, Liyuan Qiu

机构 * Hong Kong University of Science and Technology (HKUST)(香港理工大学) Hong Kong Applied Science and Technology Research Institute (ASTRI)(香港应用科技研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于扩散的算法,用于分离双层点云的内层和外层表面,特别针对具有开放边界的点云,通过提取真实内层来解决重叠表面和法线紊乱问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00583 2026-02-03 cs.CV cs.AI 79%

MAUGen: A Unified Diffusion Approach for Multi-Identity Facial Expression and AU Label Generation

MAUGen: 一种用于多身份面部表情和AU标签生成的统一扩散方法

Xiangdong Li, Ye Lou, Ao Gao, Wei Zhang, Siyang Song

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 MAUGen通过统一的扩散方法生成多身份面部表情和AU标签,提升面部动作单元识别系统的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00536 2026-02-03 cs.CV 79%

SADER: Structure-Aware Diffusion Framework with DEterministic Resampling for Multi-Temporal Remote Sensing Cloud Removal

SADER:一种结构感知扩散框架,用于多时相遥感云去除

Yifan Zhang, Qian Chen, Yi Liu, Wengen Li, Jihong Guan

机构 * College of Literature, Science, and the Arts, University of Michigan(文学、科学与艺术学院,密歇根大学) School of Computer Science and Technology, Tongji University(计算机科学与技术学院,同济大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 SADER通过结构感知扩散框架,结合时间融合和混合注意力机制,有效解决多时相遥感云去除问题,提升云去除效果和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18537 2026-02-03 cs.CV 79%

Zero-Shot Video Deraining with Video Diffusion Models

无监督视频去雨与视频扩散模型

Tuomas Varanka, Juan Luis Gonzalez, Hyeongwoo Kim, Pablo Garrido, Xu Yao

机构 * University of Oulu(奥卢大学) Flawless AI Imperial College London(伦敦帝国理工学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出了一种无需合成数据和模型微调的无监督视频去雨方法,通过预训练文本到视频扩散模型,利用注意力切换机制提升动态场景中的去雨效果。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25178 2026-02-03 cs.CV cs.AI cs.LG 79%

GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs

GHOST:诱导幻觉的多模态大语言模型图像生成

Aryan Yazdan Parast, Parsa Hosseini, Hesam Asadollahzadeh, Arshia Soltani Moakhar, Basim Azam, Soheil Feizi, Naveed Akhtar

机构 * The University of Melbourne(墨尔本大学) University of Maryland(马里兰大学)

专题命中 扩散模型 :image generation(title);diffusion(abstract);分类 cs.CV

AI总结 GHOST通过优化隐蔽令牌生成诱导幻觉的图像,评估多模态大语言模型的可靠性,并发现高幻觉成功率及可转移漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01330 2026-02-03 cs.CV 79%

Prior-Guided Residual Diffusion: Calibrated and Efficient Medical Image Segmentation

先验引导残差扩散:校准且高效的医学图像分割

Fuyou Mao, Beining Wu, Yanfeng Jiang, Han Xue, Yan Tang, Hao Zhang

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 PGRD通过先验引导残差扩散方法,在医学图像分割中实现高精度与高效校准。

Comments Withdrawn by the authors to conduct further methodological refinement and address concerns regarding the originality of the current implementation

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15078 2026-02-03 eess.IV cs.CV physics.med-ph 79%

PET Image Reconstruction Using Deep Diffusion Image Prior

基于深度扩散图像先验的PET图像重建

Fumio Hashimoto, Kuang Gong

机构 * J. Crayton Pruitt Family Department of Biomedical Engineering, University of Florida(朱·克雷顿·普瑞特家庭生物医学工程系,佛罗里达大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出基于深度扩散模型的PET图像重建方法,通过解剖先验引导和半二次分裂算法实现高效重建,适用于多种示踪剂和扫描仪类型。

Comments 11 pages, 12 figures

Journal ref IEEE Trans. Med. Imaging (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23993 2026-02-03 cs.CV cs.AI 79%

DenseFormer: Learning Dense Depth Map from Sparse Depth and Image via Conditional Diffusion Model

DenseFormer: 通过条件扩散模型学习稀疏深度和图像的密集深度图

Ming Yuan, Chuang Zhang, Lei He, Qing Xu, Jianqiang Wang

机构 * the School of Vehicle and Mobility(车辆与移动学院) Tsinghua University(清华大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DenseFormer通过条件扩散模型,结合特征提取和深度细化模块,实现从稀疏深度和图像生成密集深度图,优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.23201 2026-02-02 eess.IV cs.CV cs.LG 79%

Scale-Cascaded Diffusion Models for Super-Resolution in Medical Imaging

多尺度扩散模型用于医学影像超分辨率

Darshan Thaker, Mahmoud Mostapha, Radu Miron, Shihan Qiu, Mariappan Nadar

机构 * University of Pennsylvania(宾夕法尼亚大学) Siemens Healthineers(西门子医疗) Siemens Industry Software(西门子工业软件)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出多尺度扩散模型用于医学影像超分辨率,通过分解图像为不同频带并训练独立先验,提升重建质量并减少推理时间。

Comments Accepted at IEEE International Symposium for Biomedical Imaging (ISBI) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22744 2026-02-02 cs.CV cs.CR cs.LG 79%

Beauty and the Beast: Imperceptible Perturbations Against Diffusion-Based Face Swapping via Directional Attribute Editing

Beauty and the Beast: 面向扩散式人脸交换的不可察觉扰动防御 via 方向性属性编辑

Yilong Huang, Songze Li

机构 * Southeast University, Nanjing, China(东南大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 FaceDefense通过引入新的扩散损失和方向性属性编辑,有效提升扩散式人脸交换的防御效果和视觉不可察觉性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11181 2026-02-02 cs.CV cs.AI 79%

Multi-Stage Generative Upscaler: Reconstructing Football Broadcast Images via Diffusion Models

多阶段生成放大器:通过扩散模型重建足球转播图像

Luca Martini, Daniele Zolezzi, Saverio Iacono, Gianni Viardo Vercelli

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出多阶段生成放大框架,利用扩散模型提升足球转播图像质量,通过ControlNet和LoRA实现细节与特定元素的精准重建。

Journal ref Scientific Reports volume 16, Article number: 1882 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22468 2026-02-02 cs.CV cs.AI 79%

Training-Free Representation Guidance for Diffusion Models with a Representation Alignment Projector

无需训练的表示引导用于扩散模型的表示对齐投影器

Wenqiang Zu, Shenghao Xie, Bo Lei, Lei Ma

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Peking University(北京大学) University of Chinese Academy of Sciences(中国科学院大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出无需训练的表示引导方法,通过引入表示对齐投影器提升扩散模型的语义对齐与图像一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22275 2026-02-02 cs.CV cs.AI 79%

VMonarch: Efficient Video Diffusion Transformers with Structured Attention

VMonarch:具有结构注意力的高效视频扩散变换器

Cheng Liang, Haoxian Chen, Liang Hou, Qi Fan, Gangshan Wu, Xin Tao, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院,南京大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 VMonarch通过结构化注意力机制提升视频扩散变换器的效率,减少计算量并提高处理速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02291 2026-02-02 cs.LG cs.CV stat.ML 79%

Test-Time Anchoring for Discrete Diffusion Posterior Sampling

测试时锚定用于离散扩散后验采样

Litu Rout, Andreas Lugmayr, Yasamin Jafarian, Srivatsan Varadharajan, Constantine Caramanis, Sanjay Shakkottai, Ira Kemelmacher-Shlizerman

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Google(谷歌)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本研究提出锚定后验采样方法,通过量化期望和锚定重掩码实现离散扩散模型在后验采样中的高效和精确采样,提升图像生成和文本引导编辑的性能。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21922 2026-01-30 cs.CV 79%

Zero-Shot Video Restoration and Enhancement with Assistance of Video Diffusion Models

借助视频扩散模型的零样本视频修复与增强

Cong Cao, Huanjing Yue, Shangbin Xie, Xin Liu, Jingyu Yang

机构 * School of Electrical and Information Engineering, Tianjin University(电子信息工程学院,天津大学) Computer Vision and Pattern Recognition Laboratory, School of Engineering Science, Lappeenranta-Lahti University of Technology LUT(计算机视觉与模式识别实验室,工程科学学院,拉佩兰塔-拉赫蒂技术大学LUT)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出一种无训练的框架,利用视频扩散模型提升零样本视频修复与增强的时序一致性,通过潜在融合与后处理策略增强效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21857 2026-01-30 cs.CV 79%

Trajectory-Guided Diffusion for Foreground-Preserving Background Generation in Multi-Layer Documents

轨迹引导扩散:多层文档中保留前景的背景生成

Taewon Kang

机构 * University of Maryland at College Park(马里兰大学学院市分校)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出轨迹引导扩散方法,通过潜在空间设计实现多层文档中前景保留和背景生成的风格一致性。

Comments 47 pages, 36 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21248 2026-01-30 cs.CV 79%

NFCDS: A Plug-and-Play Noise Frequency-Controlled Diffusion Sampling Strategy for Image Restoration

NFCDS: 一种插件式噪声频率控制扩散采样策略用于图像修复

Zhen Wang, Hongyi Liu, Jianing Li, Zhihui Wei

机构 * Nanjing University of Science and Technology(南京理工大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 NFCDS通过控制噪声频率提升图像修复的保真度与感知质量,无需额外训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17868 2026-01-30 cs.CV cs.AI 79%

VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding

VidLaDA:双向扩散大型语言模型用于高效视频理解

Zhihao He, Tieyuan Chen, Kangyu Wang, Ziran Qin, Yang Shao, Chaofan Gan, Shijie Li, Zuxuan Wu, Weiyao Lin

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院) Fudan University(复旦大学) ZhongguanCun Academy(中关村学院) Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 VidLaDA通过双向注意力和MARS-Cache技术,提升视频理解效率,实现比现有方法更优的性能与速度

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11229 2026-01-30 cs.CV cs.SD 79%

REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming Distillation

REST:基于扩散模型的实时端到端流式说话人生成方法:通过ID上下文缓存和异步流式蒸馏

Haotian Wang, Yuzhe Weng, Jun Du, Haoran Xu, Xiaoyan Wu, Shan He, Bing Yin, Cong Liu, Qingfeng Liu

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 REST通过ID-Context缓存和异步流式蒸馏,实现高效实时端到端流式说话人生成。

Comments 27 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25731 2026-01-30 cs.CV 79%

LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing

LaTo:基于地标token化的扩散变换器用于精细的人脸编辑

Zhenghao Zhang, Ziying Zhang, Junchao Liao, Xiangyu Meng, Qiang Hu, Siyu Zhu, Xiaoyun Zhang, Long Qin, Weizhi Wang

机构 * Alibaba Cloud Computing(阿里云计算) Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 LaTo通过地标token化扩散变换器实现精细的人脸编辑,提升身份保持与语义一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06625 2026-01-30 cs.CV 79%

CycleDiff: Cycle Diffusion Models for Unpaired Image-to-image Translation

CycleDiff: 基于循环扩散模型的无配对图像到图像翻译

Shilong Zou, Yuhang Huang, Renjiao Yi, Chenyang Zhu, Kai Xu

机构 * College of Computer Science and Technology, National University of Defense Technology(计算机科学与技术学院,国防科技大学) College of Future Information Technology, Fudan University(未来信息技术学院,复旦大学) School of Computer, National University of Defense Technology(计算机学院,国防科技大学) Xiangjiang Laboratory(湘江实验室) Institute of AI for Industries, Chinese Academy of Sciences(产业人工智能研究院,中国科学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 CycleDiff通过联合学习扩散与翻译过程,提升跨域图像翻译的全局优化与生成质量。

Comments Accepted by IEEE TIP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07720 2026-01-30 cs.CV 79%

ACDiT: Interpolating Autoregressive Conditional Modeling and Diffusion Transformer

ACDiT:插值自回归条件建模与扩散变换器

Jinyi Hu, Shengding Hu, Yuxuan Song, Yufei Huang, Mingxuan Wang, Hao Zhou, Zhiyuan Liu, Wei-Ying Ma, Maosong Sun

机构 * Tsinghua University(清华大学) ByteDance(字节跳动)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 ACDiT通过结合自回归和扩散范式,实现连续视觉信息的灵活插值生成,优于现有自回归基线,在视觉生成任务中表现最佳。

Comments TMLR camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20857 2026-01-29 cs.CV 79%

FreeFix: Boosting 3D Gaussian Splatting via Fine-Tuning-Free Diffusion Models

FreeFix: 通过无微调扩散模型提升3D高斯散射

Hongyu Zhou, Zisen Shao, Sheng Miao, Pan Wang, Dongfeng Bai, Bingbing Liu, Yiyi Liao

机构 * Zhejiang University(浙江大学) University of Maryland, College Park(马里兰大学学院 park 分校) Huawei(华为)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 FreeFix通过无微调扩散模型提升3D高斯散射,实现高质量视图合成与强泛化能力。

Comments Our project page is at https://xdimlab.github.io/freefix

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20791 2026-01-29 cs.CV cs.AI 79%

FAIRT2V: Training-Free Debiasing for Text-to-Video Diffusion Models

FAIRT2V:无需训练的文本到视频扩散模型去偏框架

Haonan Zhong, Wei Song, Tingxu Han, Maurice Pagnucco, Jingling Xue, Yang Song

机构 * School of Computer Science and Engineering, University of New South Wales, Australia.(新南威尔士大学计算机科学与工程学院,澳大利亚) State Key Laboratory for Novel Software Technology, Nanjing University, China(南京大学新型软件技术国家重点实验室,中国)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 FAIRT2V通过无需训练的去偏框架减少文本到视频扩散模型中的性别偏见,通过中和提示嵌入和动态去噪计划实现,有效降低偏见同时保持视频质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08005 2026-01-29 cs.LG cs.CV 79%

DiffRatio: Training One-Step Diffusion Models Without Teacher Supervision

DiffRatio: 无需教师监督训练一步扩散模型

Wenlin Chen, Mingtian Zhang, Jiajun He, Zijing Ou, José Miguel Hernández-Lobato, Bernhard Schölkopf, David Barber

机构 * Bosch (China) Investment Ltd.(博世(中国)投资有限公司) University of Cambridge(剑桥大学) Imperial College London(伦敦帝国理工学院) Max Planck Institute for Intelligent Systems, Tübingen(智能系统马克斯·普朗克研究所,图宾根)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DiffRatio通过直接估计分数差来训练一步扩散模型,减少梯度偏差并提升生成质量,同时提高计算效率。

Comments 22 pages, 8 figures, 5 tables, 2 algorithms

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20564 2026-01-29 cs.CV 79%

DiffVC-RT: Towards Practical Real-Time Diffusion-based Perceptual Neural Video Compression

DiffVC-RT: 向实用的实时扩散式感知神经视频压缩迈进

Wenzhuo Ma, Zhenzhong Chen

机构 * School of Remote Sensing and Information Engineering, Wuhan University, Wuhan, China(遥感与信息工程学院,武汉大学,武汉,中国)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 DiffVC-RT通过高效架构、显式隐式一致性建模和异步并行解码实现实时扩散式视频压缩,显著提升压缩效率和质量。

Comments 17 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏