arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 70277 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70277 篇

2504.21423 2025-05-01 cs.CV 79%

Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision

Weicai Yan, Wang Lin, Zirun Guo, Ye Wang, Fangming Feng, Xiaoda Yang, Zehan Wang, Tao Jin

机构 * Zhejiang University(浙江大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepted at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21368 2025-05-01 cs.CV cs.AI 79%

Revisiting Diffusion Autoencoder Training for Image Reconstruction Quality

Pramook Khungurn, Sukit Seripanitkarn, Phonphrm Thawatdamrongkit, Supasorn Suwajanakorn

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments AI for Content Creation (AI4CC) Workshop at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21325 2025-05-01 cs.CV 79%

Text-Conditioned Diffusion Model for High-Fidelity Korean Font Generation

Abdul Sami, Avinash Kumar, Irfanullah Memon, Youngwon Jo, Muhammad Rizwan, Jaeyoung Choi

机构 * School of Computer Science and Engineering, Soongsil University(计算机科学与工程学院,顺斯大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 6 pages, 4 figures, Accepted at ICOIN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21266 2025-05-01 cs.CV 79%

CoCoDiff: Diversifying Skeleton Action Features via Coarse-Fine Text-Co-Guided Latent Diffusion

Zhifu Zhao, Hanyang Hua, Jianan Li, Shaoxin Wu, Fu Li, Yangtao Zhou, Yang Li

机构 * School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院) School of Artificial Intelligence, Xidian University(西安电子科技大学人工智能学院) Guangzhou Institute of technology, Xidian University(广州科技研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.16082 2025-05-01 cs.CV 79%

SignDiff: Diffusion Model for American Sign Language Production

Sen Fang, Chunyu Sui, Yanghao Zhou, Xuedong Zhang, Hongbin Zhong, Yapeng Tian, Chen Chen

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Camera-Ready Version; Project Page at https://signdiff.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20645 2025-04-30 cs.CV 79%

LDPoly: Latent Diffusion for Polygonal Road Outline Extraction in Large-Scale Topographic Mapping

Weiqin Jiao, Hao Cheng, George Vosselman, Claudio Persello

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20288 2025-04-30 cs.CV 79%

Image Interpolation with Score-based Riemannian Metrics of Diffusion Models

Shinnosuke Saito, Takashi Matsubara

机构 * Hokkaido University(北海道大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20111 2025-04-30 cs.CV 79%

Forging and Removing Latent-Noise Diffusion Watermarks Using a Single Image

Anubhav Jain, Yuya Kobayashi, Naoki Murata, Yuhta Takida, Takashi Shibuya, Yuki Mitsufuji, Niv Cohen, Nasir Memon, Julian Togelius

机构 * New York University(纽约大学) Sony AI(索尼人工智能) Sony Group Corporation(索尼集团)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19165 2025-04-30 cs.CV 79%

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos

Yuan Li, Ziqian Bai, Feitong Tan, Zhaopeng Cui, Sean Fanello, Yinda Zhang

机构 * State Key Lab of CAD&CG, Zhejiang University(浙江大学计算机辅助设计与图形学国家重点实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments CVPR2025; project page: https://y-u-a-n-l-i.github.io/projects/IM-Portrait/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06698 2025-04-30 cs.LG cs.CV 79%

What's in a Latent? Leveraging Diffusion Latent Space for Domain Generalization

Xavier Thomas, Deepti Ghadiyaram

机构 * Boston University(波士顿大学) Runway

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04606 2025-04-30 cs.CV cs.AI cs.CL cs.LG 79%

The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation

Aoxiong Yin, Kai Shen, Yichong Leng, Xu Tan, Xinyu Zhou, Juncheng Li, Siliang Tang

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Our code is available at https://github.com/LanDiff/LanDiff

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18565 2025-04-30 cs.CV 79%

3DEnhancer: Consistent Multi-View Diffusion for 3D Enhancement

Yihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan, Chen Change Loy

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Project page: https://yihangluo.com/projects/3DEnhancer

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19402 2025-04-29 cs.CV 79%

Boosting 3D Liver Shape Datasets with Diffusion Models and Implicit Neural Representations

Khoa Tuan Nguyen, Francesca Tozzi, Wouter Willaert, Joris Vankerschaver, Nikdokht Rashidian, Wesley De Neve

机构 * Ghent University(根特大学) Ghent University Global Campus(根特大学全球校区) Ghent University Hospital(根特大学医院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01776 2025-04-29 cs.CV cs.LG 79%

Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity

Haocheng Xi, Shuo Yang, Yilong Zhao, Chenfeng Xu, Muyang Li, Xiuyu Li, Yujun Lin, Han Cai, Jintao Zhang, Dacheng Li, Jianfei Chen, Ion Stoica, Kurt Keutzer, Song Han

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 17 pages, 11 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15309 2025-04-29 eess.IV cs.CV cs.LG 79%

Investigating the Feasibility of Patch-based Inference for Generalized Diffusion Priors in Inverse Problems for Medical Images

Saikat Roy, Mahmoud Mostapha, Radu Miron, Matt Holbrook, Mariappan Nadar

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepted at IEEE International Symposium for Biomedical Imaging (ISBI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09998 2025-04-29 eess.IV cs.AI cs.CV 79%

Self-Consistent Nested Diffusion Bridge for Accelerated MRI Reconstruction

Tao Song, Yicheng Wu, Minhao Hu, Xiangde Luo, Guoting Luo, Guotai Wang, Yi Guo, Feng Xu, Shaoting Zhang

机构 * School of Information Science and Technology, Fudan University, Shanghai, China(复旦大学信息科学与技术学院) Department of Data Science & AI, Faculty of Information Technology, Monash University, Melbourne, Australia(墨尔本大学信息科技学院数据科学与人工智能系) Nuffield Department of Clinical Neurosciences, University of Oxford, London, UK(牛津大学临床神经科学系) School of Mechanical and Electrical Engineering, University of Electronic Science and Technology of China, Chengdu, China(电子科技大学机械与电子工程学院) Department of Radiology, Sichuan Provincial People’s Hospital, Chengdu, China(四川省人民医院放射科) Department of Radiology, Zhongda Hospital, Medical School, Southeast University, Nanjing, China(东南大学医学院中大医院放射科) Shanghai AI Lab, Shanghai, China(上海人工智能实验室) Sensetime Research, Shanghai, China(商汤研究)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18808 2025-04-29 cs.CV 79%

Lifting Motion to the 3D World via 2D Diffusion

Jiaman Li, C. Karen Liu, Jiajun Wu

机构 * Stanford University(斯坦福大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments CVPR 2025 (Highlight), project page: https://lijiaman.github.io/projects/mvlift/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10619 2025-04-29 cs.CV cs.AI eess.IV 79%

Hierarchical Attention Diffusion Networks with Object Priors for Video Change Detection

Andrew Kiruluta, Eric Lundy, Andreas Lemos

机构 * School of Information, University of California, Berkeley(信息学院,加州大学伯克利分校)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15956 2025-04-29 cs.CV 79%

Anomaly Detection with Conditioned Denoising Diffusion Models

Arian Mousakhan, Thomas Brox, Jawad Tayyub

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Journal ref Proceedings of the 46th German Conference on Pattern Recognition (GCPR 2024), Lecture Notes in Computer Science, vol. 14641, Springer, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18520 2025-04-28 eess.IV cs.CV 79%

RSFR: A Coarse-to-Fine Reconstruction Framework for Diffusion Tensor Cardiac MRI with Semantic-Aware Refinement

Jiahao Huang, Fanwen Wang, Pedro F. Ferreira, Haosen Zhang, Yinzhe Wu, Zhifan Gao, Lei Zhu, Angelica I. Aviles-Rivero, Carola-Bibiane Schonlieb, Andrew D. Scott, Zohya Khalique, Maria Dwornik, Ramyah Rajakulasingam, Ranil De Silva, Dudley J. Pennell, Guang Yang, Sonia Nielles-Vallespin

机构 * Imperial College London(帝国理工学院伦敦分校) National Heart and Lung Institute(国家心肺研究所) Cardiovascular Research Centre(心血管研究中心) Royal Brompton Hospital(皇家布里蒙特医院) School of Biomedical Engineering, Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区生物医学工程学院) Robotics and Autonomous Systems Thrust & Data Science and Analytics Thrust, HKUST (GZ)(香港科技大学(广州)机器人与自主系统 thrust 及数据科学与分析 thrust) Yau Mathematical Sciences Centre, Tsinghua University(清华大学姚数学科学中心) Department of Applied Mathematics and Theoretical Physics, University of Cambridge(剑桥大学应用数学与理论物理系) School of Biomedical Engineering and Imaging Sciences, King’s College London(伦敦国王学院生物医学工程与影像科学学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17825 2025-04-28 cs.CV cs.AI 79%

Dual Prompting Image Restoration with Diffusion Transformers

Dehong Kong, Fan Li, Zhixin Wang, Jiaqi Xu, Renjing Pei, Wenbo Li, WenQi Ren

机构 * School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学网络科学与技术学院) MoE Key Laboratory of Information Technology(信息科技联合实验室) Huawei Noah’s Ark Lab(华为诺亚实验室) The Chinese University of Hong Kong(香港中文大学) Guangdong Provincial Key Laboratory of Information Security Technology(广东省信息安全技术重点实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14509 2025-04-28 cs.CV cs.AI 79%

DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning

Fulong Ye, Miao Hua, Pengze Zhang, Xinghui Li, Qichao Sun, Songtao Zhao, Qian He, Xinglong Wu

机构 * ByteDance(字节跳动)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Project: https://superhero-7.github.io/DreamID/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12633 2025-04-28 cs.RO cs.AI cs.CV cs.LG 79%

Instant Policy: In-Context Imitation Learning via Graph Diffusion

Vitalis Vosylius, Edward Johns

机构 * The Robot Learning Lab at Imperial College London(帝国理工学院伦敦分校机器人学习实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Code and videos are available on our project webpage at https://www.robot-learning.uk/instant-policy

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21669 2025-04-28 cs.CV 79%

Investigating Memorization in Video Diffusion Models

Chen Chen, Enhuai Liu, Daochang Liu, Mubarak Shah, Chang Xu

机构 * School of Computer Science, Faculty of Engineering, The University of Sydney, Australia(悉尼大学计算机科学学院,工程学院) School of Physics, Mathematics and Computing, The University of Western Australia, Australia(西澳大学物理、数学与计算学院) Center for Research in Computer Vision, University of Central Florida, USA(佛罗里达中央大学计算机视觉研究中心)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepted at DATA-FM Workshop @ ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18037 2025-04-28 cs.CV 79%

Towards Synchronous Memorizability and Generalizability with Site-Modulated Diffusion Replay for Cross-Site Continual Segmentation

Dunyuan Xu, Xi Wang, Jingyang Zhang, Pheng-Ann Heng

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments This paper is not proper to be published on arXiv, since we think some method are quite similar with one other paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17628 2025-04-25 eess.IV cs.CV 79%

Beyond Labels: Zero-Shot Diabetic Foot Ulcer Wound Segmentation with Self-attention Diffusion Models and the Potential for Text-Guided Customization

Abderrachid Hamrani, Daniela Leizaola, Renato Sousa, Jose P. Ponce, Stanley Mathis, David G. Armstrong, Anuradha Godavarty

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 12 pages, 8 figures, journal article

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17414 2025-04-25 cs.CV 79%

3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models

Min Wei, Chaohui Yu, Jingkai Zhou, Fan Wang

机构 * DAMO Academy, Alibaba Group(阿里达摩院) Hupan Lab(虎斑实验室) Zhejiang University(浙江大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Project page: 3DV-TON/" target="_blank" rel="noopener">https://2y7c3.github.io/3DV-TON/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00473 2025-04-25 cs.LG cs.CV 79%

Weak-to-Strong Diffusion with Reflection

Lichen Bai, Masashi Sugiyama, Zeke Xie

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 23 pages, 23 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04746 2025-04-25 cs.SD cs.IR cs.MM eess.AS 79%

Diff4Steer: Steerable Diffusion Prior for Generative Music Retrieval with Semantic Guidance

Xuchan Bao, Judith Yue Li, Zhong Yi Wan, Kun Su, Timo Denk, Joonseok Lee, Dima Kuzmin, Fei Sha

机构 * University of Toronto(多伦多大学) Google Research(谷歌研究) Google DeepMind(谷歌DeepMind) Seoul National University(首尔国立大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.MM

Comments NeurIPS 2024 Creative AI Track

Journal ref Proc. IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14837 2025-04-25 cs.LG cs.AI cs.CV 79%

Diffusion Models Are Real-Time Game Engines

Dani Valevski, Yaniv Leviathan, Moab Arar, Shlomi Fruchter

机构 * Google Research(谷歌研究) Google DeepMind(谷歌DeepMind) Tel Aviv University(特拉维夫大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments ICLR 2025. Project page: https://gamengen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏