arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86714 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70159 篇

2505.08190 2025-05-14 cs.CV cs.LG 83%

Unsupervised Raindrop Removal from a Single Image using Conditional Diffusion Models

Lhuqita Fazry, Valentino Vito

机构 * Faculty of Computer Science, Universitas Indonesia(计算机科学学院,印度尼西亚大学)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07906 2025-05-14 cond-mat.mtrl-sci cs.CV cs.LG 83%

Image-Guided Microstructure Optimization using Diffusion Models: Validated with Li-Mn-rich Cathode Precursors

Geunho Choi, Changhwan Lee, Jieun Kim, Insoo Ye, Keeyoung Jung, Inchul Park

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments 37 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05829 2025-05-12 cs.CV cs.LG eess.IV 83%

Accelerating Diffusion Transformer via Increment-Calibrated Caching with Channel-Aware Singular Value Decomposition

Zhiyuan Chen, Keyi Li, Yifan Jia, Le Ye, Yufei Ma

机构 * Peking University(北京大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02255 2025-05-12 cs.CV cs.AI 83%

Enhancing AI Face Realism: Cost-Efficient Quality Improvement in Distilled Diffusion Models with a Fully Synthetic Dataset

Jakub Wasala, Bartlomiej Wrzalski, Kornelia Noculak, Yuliia Tarasenko, Oliwer Krupa, Jan Kocon, Grzegorz Chodak

机构 * Department of Artificial Intelligence, Wroclaw University of Science and Technology(人工智能系,波兹南技术大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments 25th International Conference on Computational Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04729 2025-05-08 cs.SD cs.AI cs.LG cs.MM eess.AS 83%

JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models

Peike Li, Boyu Chen, Yao Yao, Yikai Wang, Allen Wang, Alex Wang

机构 * Futureverse AI Research(未来视AI研究)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.MM

Comments Github Demo Page: https://gogoduck912.github.io/Jen1-Demo-Page/

Journal ref https://ieeecai.org/2024/wp-content/pdfs/540900a773/540900a773.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03562 2025-05-07 cs.CV cs.AI 83%

Real-Time Person Image Synthesis Using a Flow Matching Model

Jiwoo Jeong, Kirok Kim, Wooju Kim, Nam-Joon Kim

机构 * Department of Industrial Engineering, Yonsei University(延世大学工业工程系) Next-Generation Semiconductor, Seoul National University(首尔国立大学下一代半导体研究所)

专题命中 扩散模型 :image synthesis(title,abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03203 2025-05-07 cs.CV 83%

PiCo: Enhancing Text-Image Alignment with Improved Noise Selection and Precise Mask Control in Diffusion Models

Chang Xie, Chenyi Zhuang, Pan Gao

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02753 2025-05-06 cs.CV 83%

Advancing Generalizable Tumor Segmentation with Anomaly-Aware Open-Vocabulary Attention Maps and Frozen Foundation Diffusion Models

Yankai Jiang, Peng Zhang, Donglin Yang, Yuan Tian, Hai Lin, Xiaosong Wang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) The University of British Columbia(不列颠哥伦比亚大学)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

Comments This paper is accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03184 2025-05-02 cs.CV 83%

Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusion

Hao Wen, Zehuan Huang, Yaohui Wang, Xinyuan Chen, Lu Sheng

机构 * School of Software, Beihang University(北航软件学院) Shanghai AI Laboratory(上海人工智能实验室) VAST(中国科学院软件研究所)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments See our project page at https://costwen.github.io/Ouroboros3D/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21292 2025-05-01 cs.CV 83%

Can We Achieve Efficient Diffusion without Self-Attention? Distilling Self-Attention into Convolutions

ZiYi Dong, Chengxing Zhou, Weijian Deng, Pengxu Wei, Xiangyang Ji, Liang Lin

机构 * Australian National University(澳大利亚国立大学) Tsinghua University(清华大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21231 2025-05-01 cs.CV cs.AI cs.LG 83%

T2ID-CAS: Diffusion Model and Class Aware Sampling to Mitigate Class Imbalance in Neck Ultrasound Anatomical Landmark Detection

Manikanta Varaganti, Amulya Vankayalapati, Nour Awad, Gregory R. Dion, Laura J. Brattain

机构 * Department of Computer Science, University of Central Florida(计算机科学系,中央佛罗里达大学) Department of Otolaryngology Head Neck Surgery, University of Cincinnati College of Medicine(耳鼻喉科与头颈外科系,辛辛那提大学医学院) Department of Internal Medicine, University of Central Florida College of Medicine(内科系,中央佛罗里达大学医学院)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments submitted to IEEE EMBC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20685 2025-04-30 cs.CV cs.HC 83%

Efficient Listener: Dyadic Facial Motion Synthesis via Action Diffusion

Zesheng Wang, Alexandre Bruckert, Patrick Le Callet, Guangtao Zhai

机构 * Nantes Université(南特大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20241 2025-04-30 cs.CV 83%

Physics-Informed Diffusion Models for SAR Ship Wake Generation from Text Prompts

Kamirul Kamirul, Odysseas Pappas, Alin Achim

机构 * Visual Information Laboratory, University of Bristol(视觉信息实验室,布里斯托尔大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments 4 pages; Submitted Machine Intelligence for GeoAnalytics and Remote Sensing (MIGARS) - 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12743 2025-04-30 cs.CV 83%

Controllable Face Synthesis with Semantic Latent Diffusion Models

Alex Ergasti, Claudio Ferrari, Tomaso Fontanini, Massimo Bertozzi, Andrea Prati

机构 * University of Parma(帕尔马大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18563 2025-04-29 cs.CR cs.AI cs.CV 83%

Backdoor Defense in Diffusion Models via Spatial Attention Unlearning

Abha Jha, Ashwath Vaithinathan Aravindan, Matthew Salaway, Atharva Sandeep Bhide, Duygu Nur Yaldiz

机构 * University of Southern California(南加州大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18204 2025-04-28 cs.CV 83%

Optimizing Multi-Round Enhanced Training in Diffusion Models for Improved Preference Understanding

Kun Li, Jianhui Wang, Yangfan He, Xinyuan Song, Ruoyu Wang, Hongyang He, Wenxin Zhang, Jiaqi Chen, Keqin Li, Sida Li, Miao Zhang, Tianyu Shi, Xueqian Wang

机构 * Xiamen University(厦门大学) University of Electronic Science and Technology of China(电子科技大学) University of Minnesota—Twin Cities(明尼苏达大学—双城分校) Emory University(埃默里大学) Tsinghua University(清华大学) University of Warwick(沃里克大学) University of the Chinese Academy of Sciences(中国科学院大学) George Washington University(乔治华盛顿大学) University of Toronto(多伦多大学) Peking University(北京大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2503.17660

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18032 2025-04-28 cs.CV 83%

Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models

Chen Chen, Daochang Liu, Mubarak Shah, Chang Xu

机构 * School of Computer Science, Faculty of Engineering, The University of Sydney, Australia(悉尼大学计算机科学学院,工程学院) School of Physics, Mathematics and Computing, The University of Western Australia, Australia(西澳大学物理、数学与计算学院) Center for Research in Computer Vision, University of Central Florida, USA(佛罗里达中央大学计算机视觉研究中心)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Accepted at CVPR 2025. Project page: https://chenchen-usyd.github.io/PRSS-Project-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21665 2025-04-28 cs.CV 83%

Exploring Local Memorization in Diffusion Models via Bright Ending Attention

Chen Chen, Daochang Liu, Mubarak Shah, Chang Xu

机构 * School of Computer Science, Faculty of Engineering, The University of Sydney, Australia(悉尼大学计算机科学学院、工程学院) School of Physics, Mathematics and Computing, The University of Western Australia, Australia(西澳大学物理、数学与计算学院) Center for Research in Computer Vision, University of Central Florida, USA(佛罗里达中央大学计算机视觉研究中心)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Accepted at ICLR 2025 (Spotlight). Project page: https://chenchen-usyd.github.io/BE-Project-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13807 2025-04-28 cs.CV 83%

Improving Consistency in Diffusion Models for Image Super-Resolution

Junhao Gu, Peng-Tao Jiang, Hao Zhang, Mi Zhou, Jinwei Chen, Wenming Yang, Bo Li

机构 * Tsinghua University(清华大学) vivo Mobile Communication Co., Ltd(vivo移动通信有限公司)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17525 2025-04-25 cs.CV 83%

Text-to-Image Alignment in Denoising-Based Models through Step Selection

Paul Grimal, Hervé Le Borgne, Olivier Ferret

机构 * Université Paris-Saclay, CEA, List(巴黎-萨克雷大学、法国原子能委员会、List)

专题命中 扩散模型 :text-to-image(title);image generation(abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17179 2025-04-25 cs.AI cs.CV cs.LG cs.RO 83%

AUTHENTICATION: Identifying Rare Failure Modes in Autonomous Vehicle Perception Systems using Adversarially Guided Diffusion Models

Mohammad Zarei, Melanie A Jutras, Eliana Evans, Mike Tan, Omid Aaramoon

机构 * MITRE Corporation(MITRE公司) Cranium Booz Allen Hamilton Virginia(Booz Allen Hamilton弗吉尼亚)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

Comments 8 pages, 10 figures. Accepted to IEEE Conference on Artificial Intelligence (CAI), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14450 2025-04-25 cs.CV 83%

Causal Disentanglement for Robust Long-tail Medical Image Generation

Weizhi Nie, Zichun Zhang, Weijie Wang, Bruno Lepri, Anan Liu, Nicu Sebe

机构 * School of Electrical and Information Engineering, Tianjin University, China(天津大学电气与信息工程学院) Department of Information Engineering and Computer Science, University of Trento, Italy(特伦托大学信息工程与计算机科学系) Fondazione Bruno Kessler, Italy(布鲁诺·克塞勒基金会)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13763 2025-04-24 cs.CV cs.AI 83%

Decoding Vision Transformers: the Diffusion Steering Lens

Ryota Takatsuki, Sonia Joseph, Ippei Fujisawa, Ryota Kanai

机构 * Araya Inc.(Araya公司) AI Alignment Network(AI对齐网络) The University of Tokyo(东京大学) Mila - Quebec AI Institute(魁北克AI研究所) McGill University(麦吉尔大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments 12 pages, 17 figures. Accepted to the CVPR 2025 Workshop on Mechanistic Interpretability for Vision (MIV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14666 2025-04-22 cs.CV 83%

Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

Kaihang Pan, Wang Lin, Zhongqi Yue, Tenglong Ao, Liyu Jia, Wei Zhao, Juncheng Li, Siliang Tang, Hanwang Zhang

机构 * Zhejiang University(浙江大学) Nanyang Technological University(新加坡国立大学) Peking University(北京大学) Huawei Singapore Research Center(华为新加坡研究中心)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Accepted by CVPR 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19902 2025-04-22 cs.CV 83%

ICE: Intrinsic Concept Extraction from a Single Image via Diffusion Models

Fernando Julio Cendra, Kai Han

机构 * Visual AI Lab, The University of Hong Kong(香港大学视觉人工智能实验室)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments CVPR 2025, Project page: https://visual-ai.github.io/ice

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14462 2025-04-22 cs.CV 83%

Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion

Jixuan He, Wanhua Li, Ye Liu, Junsik Kim, Donglai Wei, Hanspeter Pfister

机构 * Harvard University(哈佛大学) Cornell Tech(康奈尔科技) The Hong Kong Polytechnic University(香港理工大学) Boston College(波士顿大学)

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV

Comments Code is available at: https://github.com/KaKituken/affordance-aware-any. Project page at: https://kakituken.github.io/affordance-any.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14267 2025-04-22 cs.CV 83%

Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction

Li Yu, Xuanzhe Sun, Wei Zhou, Moncef Gabbouj

机构 * School of Computer Science, Nanjing University of Information Science and Technology(信息科学与技术南京大学计算机科学学院) Jiangsu Collaborative Innovation Center of Atmospheric Environment and Equipment Technology, Nanjing University of Information Science and Technology(大气环境与设备技术江苏省协同创新中心) School of Computer Science and Informatics, Cardiff University(计算机科学与信息学院卡斯尔大学) Faculty of Information Technology and Communication Sciences, Tampere University(信息科技与通讯科学学院坦佩雷大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01027 2025-04-21 cs.CV 83%

LDM-ISP: Enhancing Neural ISP for Low Light with Latent Diffusion Models

Qiang Wen, Zhefan Rao, Yazhou Xing, Qifeng Chen

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12395 2025-04-18 cs.CV 83%

InstantCharacter: Personalize Any Characters with a Scalable Diffusion Transformer Framework

Jiale Tao, Yanbing Zhang, Qixun Wang, Yiji Cheng, Haofan Wang, Xu Bai, Zhengguang Zhou, Ruihuang Li, Linqing Wang, Chunyu Wang, Qin Lin, Qinglin Lu

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Tech Report. Code is available at https://github.com/Tencent/InstantCharacter

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12100 2025-04-17 cs.CV 83%

Generalized Visual Relation Detection with Diffusion Models

Kaifeng Gao, Siqi Chen, Hanwang Zhang, Jun Xiao, Yueting Zhuang, Qianru Sun

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Under review at IEEE TCSVT. The Appendix is provided additionally

详情

展开后加载摘要…

URL PDF HTML 收藏