arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-08-26 至 2025-08-26 共收录 101 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 10 篇

2508.17760 2025-08-26 cs.CV cs.CL 92%

CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation

Mingyue Yang, Dianxi Shi, Jialu Zhou, Xinyu Wei, Leqian Li, Shaowu Yang, Chunping Qiu

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17718 2025-08-26 cs.CV cs.AI 89%

Instant Preference Alignment for Text-to-Image Diffusion Models

Yang Li, Songlin Yang, Xiaoxuan Han, Wei Wang, Jing Dong, Yueming Lyu, Ziyu Xue

机构 * New Laboratory of Pattern Recognition, CASIA(模式识别新实验室,中国科学院自动化研究所) The Hong Kong University of Science and Technology(香港科技大学) Nanjing university(南京大学) Academy of Broadcasting Science, NRTA(广播科学研究院,国家广播电视总局)

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13195 2025-08-26 cs.CV 88%

CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models

Gaoyang Zhang, Bingtao Fu, Qingnan Fan, Qi Zhang, Runxing Liu, Hong Gu, Huaqi Zhang, Xinguo Liu

机构 * Zhejiang University(浙江大学) vivo Ant Group(蚂蚁集团)

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);分类 cs.CV

Comments 21 pages, 12 figures. Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13211 2025-08-26 cs.CV cs.AI 88%

MedLoRD: A Medical Low-Resource Diffusion Model for High-Resolution 3D CT Image Synthesis

Marvin Seyfarth, Salman Ul Hassan Dar, Isabelle Ayx, Matthias Alexander Fink, Stefan O. Schoenberg, Hans-Ulrich Kauczor, Sandy Engelhardt

机构 * Institute for Artificial Intelligence in Cardiovascular Medicine(心血管医学人工智能研究所) Heidelberg University Hospital(海德堡大学医院) AI Health Innovation Cluster (AIH)(人工智能健康创新集群) Heidelberg Faculty of Medicine(海德堡医学院) German Centre for Cardiovascular Research (DZHK)(德国心脏病研究中心) University Medical Center Mannheim(曼海姆大学医学中心) Clinic for Diagnostic and Interventional Radiology(诊断与介入放射科诊所)

专题命中 文生图 :diffusion(title,abstract);image synthesis(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17472 2025-08-26 cs.CV 86%

T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation

Kaiyue Sun, Rongyao Fang, Chengqi Duan, Xian Liu, Xihui Liu

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 文生图 :text-to-image(title,abstract);image generation(title);分类 cs.CV

Comments Code: https://github.com/KaiyueSun98/T2I-ReasonBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16752 2025-08-26 cs.CV 85%

A Framework for Benchmarking Fairness-Utility Trade-offs in Text-to-Image Models via Pareto Frontiers

Marco N. Bochernitsan, Rodrigo C. Barros, Lucas S. Kupssinskü

专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14174 2025-08-26 cs.HC 78%

Steering Large Text-to-Image Model for Abstract Art Synthesis: Preference-based Prompt Optimization and Visualization

Aven-Le Zhou, Wei Wu, Yu-Ao Wang, Kang Zhang

专题命中 文生图 :text-to-image(title,abstract)

Comments arXiv admin note: text overlap with arXiv:2402.06389

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08505 2025-08-26 cs.CV eess.IV 77%

Yuan: Yielding Unblemished Aesthetics Through A Unified Network for Visual Imperfections Removal in Generated Images

Zhenyu Yu, Chee Seng Chan

专题命中 文生图 :text-to-image(abstract);inpainting(abstract);image synthesis(abstract);分类 cs.CV

Journal ref AAAI 2025, Vol. 39, No. 9, pp. 9716-9724

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03001 2025-08-26 cs.CV cs.MM 73%

One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning

Hao Sun, Yu Song, Jiaqing Liu, Jihong Hu, Yen-Wei Chen, Lanfen Lin

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) College of Information Science and Engineering, Ritsumeikan University(立命馆大学信息科学与工程学院)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12698 2025-08-26 eess.IV cs.CV 70%

Pixel Perfect MegaMed: A Megapixel-Scale Vision-Language Foundation Model for Generating High Resolution Medical Images

Zahra TehraniNasab, Hujun Ni, Amar Kumar, Tal Arbel

机构 * McGill University(麦吉尔大学) MILA-Quebec AI Institute(魁北克AI研究所)

专题命中 文生图 :image generation(abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 图像编辑 3 篇

2412.04715 2025-08-26 cs.CV 88%

Addressing Text Embedding Leakage in Diffusion-based Image Editing

Sunung Mun, Jinhwan Nam, Sunghyun Cho, Jungseul Ok

机构 * Graduate School of AI, POSTECH(POSTECH人工智能研究生院) Dept. of CSE, POSTECH(POSTECH计算机科学与工程系)

专题命中 图像编辑 :diffusion(title,abstract);image editing(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17435 2025-08-26 cs.CV 83%

An LLM-LVLM Driven Agent for Iterative and Fine-Grained Image Editing

Zihan Liang, Jiahao Sun, Haoran Ma

机构 * Kunming University of Science and Technology(昆明理工大学)

专题命中 图像编辑 :image editing(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17315 2025-08-26 cs.CV 57%

Defending Deepfake via Texture Feature Perturbation

Xiao Zhang, Changfang Chen, Tianyi Wang

机构 * Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Qilu University of Technology (Shandong Academy of Sciences)(计算机网络与信息安全重点实验室,教育部,齐鲁工业大学(山东省科学院)) Shandong Artificial Intelligence Institute, Qilu University of Technology (Shandong Academy of Sciences)(山东省人工智能研究院,齐鲁工业大学(山东省科学院)) School of Computing, National University of Singapore(新加坡国立大学计算机学院)

专题命中 图像编辑 :image editing(abstract);分类 cs.CV

Comments Accepted to IEEE SMC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 扩散模型 69 篇

2508.15093 2025-08-26 cs.CV 85%

CurveFlow: Curvature-Guided Flow Matching for Image Generation

Yan Luo, Drake Du, Hao Huang, Yi Fang, Mengyu Wang

专题命中 扩散模型 :image generation(title,abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16827 2025-08-26 cs.GR cs.CV cs.LG 84%

Beyond Blur: A Fluid Perspective on Generative Diffusion Models

Grzegorz Gruszczynski, Jakub Meixner, Michal Jan Wlodarczyk, Przemyslaw Musialski

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV、cs.GR

Comments ICCV 2025 main conference, 8 pages paper, 20 pages appendix, 24 figures, supplementary pseudocode in appendix, https://iccv.thecvf.com/virtual/2025/poster/1176

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18235 2025-08-26 cs.CV 83%

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation

Ashwath Vaithinathan Aravindan, Abha Jha, Matthew Salaway, Atharva Sandeep Bhide, Duygu Nur Yaldiz

机构 * University of Southern California(南加州大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17844 2025-08-26 cs.CV cs.LG 83%

Diffusion-Based Data Augmentation for Medical Image Segmentation

Maham Nazir, Muhammad Aqeel, Francesco Setti

机构 * School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) Dept. of Engineering for Innovation Medicine, University of Verona(威尼斯大学创新医学工程系)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(abstract);分类 cs.CV

Comments Accepted to CVAMD Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17614 2025-08-26 cs.CV 83%

JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on

Aowen Wang, Wei Li, Hao Luo, Mengxing Ao, Chenyu Zhu, Xinyang Li, Fan Wang

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab(汇安实验室) Zhejiang University(浙江大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17045 2025-08-26 cs.CV 83%

Styleclone: Face Stylization with Diffusion Based Data Augmentation

Neeraj Matiyali, Siddharth Srivastava, Gaurav Sharma

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16212 2025-08-26 cs.CV cs.AI cs.LG 83%

OmniCache: A Trajectory-Oriented Global Perspective on Training-Free Cache Reuse for Diffusion Transformer Models

Huanpeng Chu, Wei Wu, Guanyu Fen, Yutao Zhang

机构 * Zhipu AI(智谱AI)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11688 2025-08-26 cs.CR cs.AI cs.MM 83%

Watermarking Visual Concepts for Diffusion Models

Liangqi Lei, Keke Gai, Jing Yu, Liehuang Zhu, Qi Wu

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17465 2025-08-26 cs.CY cs.AI 82%

Bias Amplification in Stable Diffusion's Representation of Stigma Through Skin Tones and Their Homogeneity

Kyra Wilson, Sourojit Ghosh, Aylin Caliskan

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract)

Comments Published in Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society; code available at https://github.com/kyrawilson/Image-Generation-Bias

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18095 2025-08-26 cs.CV cs.LG 79%

Incorporating Pre-trained Diffusion Models in Solving the Schrödinger Bridge Problem

Zhicong Tang, Tiankai Hang, Shuyang Gu, Dong Chen, Baining Guo

机构 * Tsinghua University(清华大学) Southeast University(东南大学) University of Science and Technology of China(中国科学技术大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17326 2025-08-26 eess.IV cs.CV 79%

Semantic Diffusion Posterior Sampling for Cardiac Ultrasound Dehazing

Tristan S. W. Stevens, Oisín Nolan, Ruud J. G. van Sloun

机构 * Eindhoven University of Technology(埃因霍温理工大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 10 pages, 4 figures, MICCAI challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17299 2025-08-26 cs.CV 79%

FoundDiff: Foundational Diffusion Model for Generalizable Low-Dose CT Denoising

Zhihao Chen, Qi Gao, Zilong Li, Junping Zhang, Yi Zhang, Jun Zhao, Hongming Shan

机构 * Institute of Science and Technology for Brain-inspired Intelligence and MOE Frontiers Center for Brain Science, Fudan University(脑启发智能科学技术研究所和脑科学前沿中心,复旦大学) Shanghai Key Lab of Intelligent Information Processing, School of Computer Science, Fudan University(智能信息处理上海市重点实验室,复旦大学计算机科学学院) School of Cyber Science and Engineering, Sichuan University(四川大学网络科学与工程学院) School of Biomedical Engineering, Shanghai Jiao Tong University(上海交通大学生物医学工程学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17050 2025-08-26 cs.CV 79%

PVNet: Point-Voxel Interaction LiDAR Scene Upsampling Via Diffusion Models

Xianjing Cheng, Lintai Wu, Zuowen Wang, Junhui Hou, Jie Wen, Yong Xu

机构 * Telecom Guizhou Branch(贵州电信支局) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) College of Engineering, Huaqiao University(华侨大学工程学院) Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) Bio-Computing Research Center, Shenzhen Graduate School, Harbin Institute of Technology(哈尔滨工业大学深圳研究生院生物计算研究中心) Key Laboratory of Network Oriented Intelligent Computation(网络导向智能计算重点实验室)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17017 2025-08-26 cs.CV 79%

Dual Orthogonal Guidance for Robust Diffusion-based Handwritten Text Generation

Konstantina Nikolaidou, George Retsinas, Giorgos Sfikas, Silvia Cascianelli, Rita Cucchiara, Marcus Liwicki

机构 * Luleå University of Technology(卢莱大学) National Technical University of Athens(雅典技术大学) University of West Attica(西阿提卡大学) University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 10 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16956 2025-08-26 cs.CV 79%

RPD-Diff: Region-Adaptive Physics-Guided Diffusion Model for Visibility Enhancement under Dense and Non-Uniform Haze

Ruicheng Zhang, Puxin Yan, Zeyu Zhang, Yicheng Chang, Hongyi Chen, Zhi Jin

机构 * School of Intelligent Systems Engineering, Shenzhen Campus of Sun Yat-sen University(中山大学智能系统工程学院) Guangdong Provincial Key Laboratory of Fire Science and Intelligent Emergency Technology(广东省火灾科学与智能应急技术重点实验室) The Australian National University(澳大利亚国立大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16930 2025-08-26 eess.AS cs.CV cs.SD 79%

HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Sizhe Shan, Qiulin Li, Yutao Cui, Miles Yang, Yuehai Wang, Qun Yang, Jin Zhou, Zhao Zhong

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07903 2025-08-26 eess.IV cs.AI cs.CV 79%

Diffusing the Blind Spot: Uterine MRI Synthesis with Diffusion Models

Johanna P. Müller, Anika Knupfer, Pedro Blöss, Edoardo Berardi Vittur, Bernhard Kainz, Jana Hutter

机构 * Imperial College London, London, UK(伦敦帝国理工学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepted at MICCAI CAPI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏