arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-08-12 至 2025-08-12 共收录 108 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 8 篇

2502.20667 2025-08-12 cs.CV cs.AI cs.LG 90%

Advancing AI-Powered Medical Image Synthesis: Insights from MedVQA-GI Challenge Using CLIP, Fine-Tuned Stable Diffusion, and Dream-Booth + LoRA

Ojonugwa Oluwafemi Ejiga Peter, Md Mahmudur Rahman, Fahmi Khalifa

机构 * Department of Computer Science, SCMNS School, Morgan State University, Baltimore, Maryland 21251, USA Electrical \& Computer Engineering Dept., School of Engineering, Morgan State University, Baltimore, Maryland 21251, USA

专题命中 文生图 :diffusion(title,abstract);image synthesis(title,abstract);image generation(abstract);text-to-image(abstract)

Journal ref Conference and Labs of the Evaluation Forum (CLEF) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05658 2025-08-12 cs.CR cs.CV cs.MM 86%

Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards

Song Yan, Hui Wei, Jinlong Fei, Guoliang Yang, Zhengyu Zhao, Zheng Wang

机构 * Information Engineering University Zhengzhou China School of Computer Science, \ University Wuhan China Xi’an Jiaotong University Xi’an China Wuhan University Wuhan China Information Engineering University School of Computer Science, \ University Xi’an Jiaotong University Wuhan University

专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract);inpainting(abstract);分类 cs.CV、cs.MM

Comments This paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06924 2025-08-12 cs.CV 85%

AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning

Shihao Yuan, Yahui Liu, Yang Yue, Jingyuan Zhang, Wangmeng Zuo, Qi Wang, Fuzheng Zhang, Guorui Zhou

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);image synthesis(abstract);分类 cs.CV

Comments 27 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06916 2025-08-12 cs.CV 85%

Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing

Shichao Ma, Yunhe Guo, Jiahao Su, Qihe Huang, Zhengyang Zhou, Yang Wang

专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);image editing(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07897 2025-08-12 cs.CV cs.AI 74%

NeeCo: Image Synthesis of Novel Instrument States Based on Dynamic and Deformable 3D Gaussian Reconstruction

Tianle Zeng, Junlei Hu, Gerardo Loza Galindo, Sharib Ali, Duygu Sarikaya, Pietro Valdastri, Dominic Jones

机构 * STORM Lab UK, School of Electronic and Electrical Engineering, University of Leeds(STORM实验室(英国)、电子与电气工程学院、利兹大学)

专题命中 文生图 :image synthesis(title);分类 cs.CV

Comments 13 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21839 2025-08-12 cs.CV cs.CL 70%

GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles

Mengyi Shan, Brian Curless, Ira Kemelmacher-Shlizerman, Steve Seitz

机构 * University of Washington(华盛顿大学)

专题命中 文生图 :text-to-image(abstract);image editing(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08123 2025-08-12 eess.IV cs.CV 57%

A Physics-Driven Neural Network with Parameter Embedding for Generating Quantitative MR Maps from Weighted Images

Lingjing Chen, Chengxiu Zhang, Yinqiao Yi, Yida Wang, Yang Song, Xu Yan, Shengfang Xu, Dalin Zhu, Mengqiu Cao, Yan Zhou, Chenglong Wang, Guang Yang

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06715 2025-08-12 cs.CV 57%

Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video

Jixuan He, Chieh Hubert Lin, Lu Qi, Ming-Hsuan Yang

机构 * Cornell Tech(康奈尔科技) University of California, Merced(加州大学默塞德分校) Wuhan University(武汉大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 图像编辑 1 篇

2504.10434 2025-08-12 cs.CV 87%

Anchor Token Matching: Implicit Structure Locking for Training-free AR Image Editing

Taihang Hu, Linxuan Li, Kai Wang, Yaxing Wang, Jian Yang, Ming-Ming Cheng

机构 * VCIP, College of Computer Science, Nankai University(VCIP,计算机科学学院,南开大学) NKIARI, Shenzhen Futian(NKIARI,深圳福田) Computer Vision Center, Universitat Autònoma de Barcelona(计算机视觉中心,巴塞罗那自治大学) City University of Hong Kong (Dongguan)(香港城市大学(东莞))

专题命中 图像编辑 :image editing(title,abstract);image generation(abstract);text-to-image(abstract);diffusion(abstract)

Comments Accepted by ICCV2025. Code will be released in https://github.com/hutaiHang/ATM

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 扩散模型 78 篇

2505.22792 2025-08-12 cs.CV 92%

Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization

Yuxi Zhang, Yueting Li, Xinyu Du, Sibo Wang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of California, Berkeley(加州大学伯克利分校)

专题命中 扩散模型 :image generation(title,abstract);text-to-image(title,abstract);diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07519 2025-08-12 cs.CV 88%

Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing

Joonghyuk Shin, Alchan Hwang, Yujin Kim, Daneul Kim, Jaesik Park

机构 * Seoul National University(首尔国立大学)

专题命中 扩散模型 :diffusion(title,abstract);image editing(title,abstract);分类 cs.CV

Comments ICCV 2025. Project webpage: https://joonghyuk.com/exploring-mmdit-web/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16726 2025-08-12 cs.CV cs.LG 87%

EDiT: Efficient Diffusion Transformers with Linear Compressed Attention

Philipp Becker, Abhinav Mehrotra, Ruchika Chavhan, Malcolm Chadwick, Luca Morreale, Mehdi Noroozi, Alberto Gil Ramos, Sourav Bhattacharya

机构 * Samsung, AI Center Cambridge(三星人工智能中心剑桥)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06134 2025-08-12 cs.CV 87%

X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation

Jian Ma, Qirong Peng, Xu Guo, Chen Chen, Haonan Lu, Zhenyu Yang

机构 * OPPO AI Center(OPPO人工智能中心) Tsinghua University(清华大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);image editing(abstract)

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07747 2025-08-12 cs.CV 83%

Grouped Speculative Decoding for Autoregressive Image Generation

Junhyuk So, Juncheol Shin, Hyunho Kook, Eunhyeok Park

机构 * Department of Computer Science and Engineering, POSTECH(计算机科学与工程系,POSTECH) Graduate School of Artificial Intelligence, POSTECH(人工智能研究生院,POSTECH)

专题命中 扩散模型 :image generation(title,abstract);diffusion(abstract);分类 cs.CV

Comments Accepted to the ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07183 2025-08-12 cs.HC cs.AI cs.LG cs.MM 83%

Explainability-in-Action: Enabling Expressive Manipulation and Tacit Understanding by Bending Diffusion Models in ComfyUI

Ahmed M. Abuzuraiq, Philippe Pasquier

机构 * School of Interactive Arts and Technology(交互艺术与技术学院) Simon Fraser University(西蒙弗雷泽大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.MM

Comments In Proceedings of Explainable AI for the Arts Workshop 2025 (XAIxArts 2025) arXiv:2406.14485

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.04832 2025-08-12 cs.LG cs.AI math.OC 82%

Reward-Directed Score-Based Diffusion Models via q-Learning

Xuefeng Gao, Jiale Zha, Xun Yu Zhou

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07413 2025-08-12 cs.CV 81%

CLUE: Leveraging Low-Rank Adaptation to Capture Latent Uncovered Evidence for Image Forgery Localization

Youqi Wang, Shunquan Tan, Rongxuan Peng, Bin Li, Jiwu Huang

专题命中 扩散模型 :text-to-image(abstract);diffusion(abstract);image editing(abstract);image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06923 2025-08-12 cs.CV cs.AI 80%

From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers

Jiacheng Liu, Chang Zou, Yuanhuiyi Lyu, Junjie Chen, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shandong University(山东大学) University of Electronic Science and Technology of China(电子科技大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 15 pages, 14 figures; Accepted by ICCV2025; Mainly focus on feature caching for diffusion transformers acceleration

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07700 2025-08-12 cs.CV 79%

Make Your MoVe: Make Your 3D Contents by Adapting Multi-View Diffusion Models to External Editing

Weitao Wang, Haoran Xu, Jun Meng, Haoqian Wang

机构 * Tsinghua University(清华大学) Zhejiang University(浙江大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07682 2025-08-12 eess.IV cs.CV 79%

DiffVC-OSD: One-Step Diffusion-based Perceptual Neural Video Compression Framework

Wenzhuo Ma, Zhenzhong Chen

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07557 2025-08-12 cs.CV 79%

Splat4D: Diffusion-Enhanced 4D Gaussian Splatting for Temporally and Spatially Consistent Content Creation

Minghao Yin, Yukang Cao, Songyou Peng, Kai Han

机构 * The University of Hong Kong(香港大学) Nanyang Technological University(南洋理工大学) Google DeepMind(谷歌DeepMind)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07346 2025-08-12 cs.CV 79%

SODiff: Semantic-Oriented Diffusion Model for JPEG Compression Artifacts Removal

Tingyu Yang, Jue Gong, Jinpei Guo, Wenbo Li, Yong Guo, Yulun Zhang

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 7 pages, 5 figures. The code will be available at \url{https://github.com/frakenation/SODiff}

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07162 2025-08-12 cs.CV 79%

CoopDiff: Anticipating 3D Human-object Interactions via Contact-consistent Decoupled Diffusion

Xiaotong Lin, Tianming Liang, Jian-Fang Hu, Kun-Yu Lin, Yulei Kang, Chunwei Tian, Jianhuang Lai, Wei-Shi Zheng

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07146 2025-08-12 cs.CV cs.AI 79%

Intention-Aware Diffusion Model for Pedestrian Trajectory Prediction

Yu Liu, Zhijie Liu, Xiao Ren, You-Fu Li, He Kong

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07006 2025-08-12 eess.IV cs.CV 79%

Spatio-Temporal Conditional Diffusion Models for Forecasting Future Multiple Sclerosis Lesion Masks Conditioned on Treatments

Gian Mario Favero, Ge Ya Luo, Nima Fathi, Justin Szeto, Douglas L. Arnold, Brennan Nichyporuk, Chris Pal, Tal Arbel

机构 * McGill University(麦吉尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepted to MICCAI 2025 (LMID Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03295 2025-08-12 cs.CV 79%

CPKD: Clinical Prior Knowledge-Constrained Diffusion Models for Surgical Phase Recognition in Endoscopic Submucosal Dissection

Xiangning Zhang, Jinnan Chen, Qingwei Zhang, Yaqi Wang, Chengfeng Zhou, Xiaobo Li, Dahong Qian

机构 * School of Biomedical Engineering(生物医学工程学院) Division of Gastroenterology and Hepatology, Shanghai Institute of Digestive Disease, NHC Key Laboratory of Digestive Diseases, Renji Hospital(消化内科与肝病科、上海消化疾病研究所、国家消化疾病临床医学研究中心、仁济医院) College of Media Engineering(媒体工程学院) Aier Institute of Digital Ophthalmology and Visual Science(数字眼科与视觉科学研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19914 2025-08-12 cs.CV 79%

Learning 3D Object Spatial Relationships from Pre-trained 2D Diffusion Models

Sangwon Baik, Hyeonwoo Kim, Hanbyul Joo

机构 * Seoul National University(首尔国立大学) RLWRLD

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Project Page: https://tlb-miss.github.io/oor/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08333 2025-08-12 cs.CV 79%

DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models

Hyeonwoo Kim, Sangwon Baik, Hanbyul Joo

机构 * Seoul National University(首尔国立大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Project Page: https://snuvclab.github.io/david/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01440 2025-08-12 cs.CV 79%

BadPatch: Diffusion-Based Generation of Physical Adversarial Patches

Zhixiang Wang, Xingjun Ma, Yu-Gang Jiang

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理关键实验室,计算机科学学院,复旦大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Code available at: https://github.com/Wwangb/BadPatch

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12777 2025-08-12 cs.CV cs.CL cs.CR cs.LG 79%

Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts

Hongcheng Gao, Tianyu Pang, Chao Du, Taihang Hu, Zhijie Deng, Min Lin

机构 * Sea AI Lab, Singapore(新加坡Sea AI实验室) University of Chinese Academy of Sciences(中国科学院大学) Shanghai Jiao Tong University(上海交通大学) Nankai University(南开大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏