arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 70159 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 70159 篇

2508.05755 2025-08-11 cs.CV cs.AI 83%

UnGuide: Learning to Forget with LoRA-Guided Diffusion Models

Agnieszka Polowczyk, Alicja Polowczyk, Dawid Malarz, Artur Kasymov, Marcin Mazur, Jacek Tabor, Przemysław Spurek

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15877 2025-08-08 cs.CV 83%

Repurposing 2D Diffusion Models with Gaussian Atlas for 3D Generation

Tiange Xiang, Kai Li, Chengjiang Long, Christian Häne, Peihong Guo, Scott Delp, Ehsan Adeli, Li Fei-Fei

机构 * Stanford University(斯坦福大学) Meta Reality Labs(Meta现实实验室)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14404 2025-08-06 cs.CV cs.AI 83%

Causally Steered Diffusion for Automated Video Counterfactual Generation

Nikos Spyrou, Athanasios Vlontzos, Paraskevas Pegios, Thomas Melistas, Nefeli Gkouti, Yannis Panagakis, Giorgos Papanastasiou, Sotirios A. Tsaftaris

机构 * Spotify, UK(英国Spotify)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00482 2025-08-06 cs.CV cs.AI 83%

JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers

Kwon Byung-Ki, Qi Dai, Lee Hyoseok, Chong Luo, Tae-Hyun Oh

机构 * POSTECH Microsoft Research Asia(微软亚洲研究院) KAIST(韩国科学技术院)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Accepted to IEEE/CVF International Conference on Computer Vision (ICCV) 2025. Project page: https://byungki-k.github.io/JointDiT/ Code: https://github.com/kaist-ami/JointDiT

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14432 2025-08-06 cs.CV eess.IV 83%

IntroStyle: Training-Free Introspective Style Attribution using Diffusion Features

Anand Kumar, Jiteng Mu, Nuno Vasconcelos

机构 * University of California, San Diego(加州大学圣地亚哥分校)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments 17 pages, 16 figures

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15580 2025-08-05 cs.CV 83%

TKG-DM: Training-free Chroma Key Content Generation Diffusion Model

Ryugo Morita, Stanislav Frolov, Brian Bernhard Moser, Takahiro Shirakawa, Ko Watanabe, Andreas Dengel, Jinjia Zhou

机构 * Faculty of Science and Engineering(科学与工程学部) RPTU Kaiserslautern-Landau & DFKI GmbH(科隆-兰道大学与DFKI GmbH)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Accepted to CVPR2025(Highlight). Code at: https://github.com/ryugo417/TKG-DM

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18092 2025-08-05 cs.CV cs.AI cs.RO 83%

DiffSSC: Semantic LiDAR Scan Completion using Denoising Diffusion Probabilistic Models

Helin Cao, Sven Behnke

机构 * Autonomous Intelligent Systems group, Computer Science Institute VI – Intelligent Systems and Robotics – and the Center for Robotics and the Lamarr Institute for Machine Learning and Artificial Intelligence, University of Bonn, Germany(自主智能系统组,计算机科学研究所VI——智能系统与机器人——和机器人中心以及拉马尔人工智能与机器学习研究所,波恩大学,德国)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2025), Hangzhou, China, Oct 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01130 2025-08-05 cs.CV 83%

Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models

Bicheng Xu, Qi Yan, Renjie Liao, Lele Wang, Leonid Sigal

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute for AI(人工智能向量研究所) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00438 2025-08-04 eess.IV cs.CV 83%

Diffusion-Based User-Guided Data Augmentation for Coronary Stenosis Detection

Sumin Seo, In Kyu Lee, Hyun-Woo Kim, Jaesik Min, Chung-Hwan Jung

机构 * Medipixel, Inc.(Medipixel公司) University of California San Diego(加州大学圣地亚哥分校)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

Comments Accepted at MICCAI 2025. Dataset available at https://github.com/medipixel/DiGDA

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00413 2025-08-04 cs.CV cs.AI 83%

DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space

Junyu Chen, Dongyun Zou, Wenkun He, Junsong Chen, Enze Xie, Song Han, Han Cai

机构 * NVIDIA

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05846 2025-08-01 cs.CR cs.CV 83%

An Inversion-based Measure of Memorization for Diffusion Models

Zhe Ma, Qingming Li, Xuhong Zhang, Tianyu Du, Ruixiao Lin, Zonghui Wang, Shouling Ji, Wenzhi Chen

机构 * Zhejiang University(浙江大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01654 2025-08-01 cs.LG cs.CV stat.ML 83%

Insights into Closed-form IPM-GAN Discriminator Guidance for Diffusion Modeling

Aadithya Srikanth, Siddarth Asokan, Nishanth Shetty, Chandra Sekhar Seelamantula

机构 * Microsoft Research #9 VIGYAN(微软研究院)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00983 2025-07-31 eess.IV cs.CV 83%

DMCIE: Diffusion Model with Concatenation of Inputs and Errors to Improve the Accuracy of the Segmentation of Brain Tumors in MRI Images

Sara Yavari, Rahul Nitin Pandya, Jacob Furst

机构 * School of Computing, DePaul University(计算学院,德保罗大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21195 2025-07-30 cs.CR cs.AI cs.MM 83%

MaXsive: High-Capacity and Robust Training-Free Generative Image Watermarking in Diffusion Models

Po-Yuan Mao, Cheng-Chang Tsai, Chun-Shien Lu

机构 * IIS, Academia Sinica(中国台湾“中央研究院”资讯研究所)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17226 2025-07-29 cs.CV 83%

DDB: Diffusion Driven Balancing to Address Spurious Correlations

Aryan Yazdan Parast, Basim Azam, Naveed Akhtar

机构 * The University of Melbourne(墨尔本大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05367 2025-07-24 cs.CV 83%

Text2Stereo: Repurposing Stable Diffusion for Stereo Generation with Consistency Rewards

Aakash Garg, Libing Zeng, Andrii Tsarov, Nima Khademi Kalantari

机构 * Texas A&M University(德克萨斯A&M大学) Leia Inc(Leia公司)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07986 2025-07-24 cs.CV 83%

Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers

Zhengyao Lv, Tianlin Pan, Chenyang Si, Zhaoxi Chen, Wangmeng Zuo, Ziwei Liu, Kwan-Yee K. Wong

机构 * The University of Hong Kong(香港大学) Nanjing University(南京大学) University of Chinese Academy of Sciences(中国科学院大学) Nanyang Technological University(南洋理工大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Accepted by ICCV 2025; Project Page: https://vchitect.github.io/TACA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16579 2025-07-23 eess.IV cs.AI cs.CV 83%

Pyramid Hierarchical Masked Diffusion Model for Imaging Synthesis

Xiaojiao Xiao, Qinmin Vivian Hu, Guanghui Wang

机构 * Department of Computer Science(计算机科学系) Toronto Metropolitan University(多伦多 Metropolitan 大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01738 2025-07-23 cs.CV cs.AI 83%

VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models

Kailai Feng, Yabo Zhang, Haodong Yu, Zhilong Ji, Jinfeng Bai, Hongzhi Zhang, Wangmeng Zuo

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments https://github.com/Carlofkl/VitaGlyph

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15361 2025-07-22 eess.IV cs.AI cs.CV 83%

Latent Space Synergy: Text-Guided Data Augmentation for Direct Diffusion Biomedical Segmentation

Muhammad Aqeel, Maham Nazir, Zanxi Ruan, Francesco Setti

机构 * Dept. of Engineering for Innovation Medicine, University of Verona(创新医学工程系,威尼斯大学) Department of Computer Science, Beihang University(计算机科学系,北航大学)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

Comments Accepted to CVGMMI Workshop at ICIAP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15216 2025-07-22 cs.CV 83%

Improving Joint Embedding Predictive Architecture with Diffusion Noise

Yuping Qiu, Rui Zhu, Ying-cong Chen

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14797 2025-07-22 cs.CV 83%

Distilling Parallel Gradients for Fast ODE Solvers of Diffusion Models

Beier Zhu, Ruoyu Wang, Tong Zhao, Hanwang Zhang, Chi Zhang

机构 * Nanyang Technological University(南洋理工大学) Westlake University(西湖大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

Comments To appear in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14575 2025-07-22 cs.CV cs.AI 83%

Benchmarking GANs, Diffusion Models, and Flow Matching for T1w-to-T2w MRI Translation

Andrea Moschetto, Lemuel Puglisi, Alec Sargood, Pierluigi Dell'Acqua, Francesco Guarnera, Sebastiano Battiato, Daniele Ravì

机构 * Università degli Studi di Catania(卡塔尼亚大学) Università degli Studi di Messina(梅斯纳大学) University College London(伦敦大学学院)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12170 2025-07-21 cs.RO cs.CV 83%

DiffAD: A Unified Diffusion Modeling Approach for Autonomous Driving

Tao Wang, Cong Zhang, Xingguang Qu, Kun Li, Weiwei Liu, Chang Huang

机构 * Carizon Beihang University(北航大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments 8 pages, 6 figures; Code released

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06848 2025-07-21 cs.LG cs.CL cs.CV 83%

A General Framework for Inference-time Scaling and Steering of Diffusion Models

Raghav Singhal, Zachary Horvitz, Ryan Teehan, Mengye Ren, Zhou Yu, Kathleen McKeown, Rajesh Ranganath

机构 * Department of Computer Science, New York University(纽约大学计算机科学系) Columbia University(哥伦比亚大学) Center for Data Science, New York University(纽约大学数据科学中心)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12933 2025-07-18 cs.CV cs.AI cs.LG 83%

DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization

Dongyeun Lee, Jiwan Hur, Hyounguk Shon, Jae Young Lee, Junmo Kim

机构 * KAIST(韩国科学技术院)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03558 2025-07-18 cs.CV 83%

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Zehuan Huang, Yuan-Chen Guo, Xingqiao An, Yunhan Yang, Yangguang Li, Zi-Xin Zou, Ding Liang, Xihui Liu, Yan-Pei Cao, Lu Sheng

机构 * Beihang University(北京航空航天大学) VAST Tsinghua University(清华大学) The University of Hong Kong(香港大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Project page: https://huanngzh.github.io/MIDI-Page/

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 2025, pp. 23646 - 23657

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09502 2025-07-18 cs.LG cs.CV 83%

Golden Noise for Diffusion Models: A Learning Framework

Zikai Zhou, Shitong Shao, Lichen Bai, Shufei Zhang, Zhiqiang Xu, Bo Han, Zeke Xie

机构 * HKUST-GZ(香港科技大学-广州) Shanghai AI Lab(上海人工智能实验室) MBZUAI(穆罕默德·本·拉希德人工智能学院) HKBU(香港都会大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10745 2025-07-17 cs.CV 83%

Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition

Jeonghyeok Do, Munchurl Kim

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments ICCV 2025 (camera-ready version). Please visit our project page at https://kaist-viclab.github.io/TDSM_site/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04692 2025-07-15 cs.CV 83%

Structure-Guided Diffusion Models for High-Fidelity Portrait Shadow Removal

Wanchang Yu, Qing Zhang, Rongjia Zheng, Wei-Shi Zheng

机构 * School of Computer Science and Engineering, Sun Yat-sen University, China(中山大学计算机科学与工程学院) Key Laboratory of Machine Intelligence and Advanced Computing, Ministry of Education, China(教育部人工智能与先进计算重点实验室)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏