arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-10-14 至 2025-10-14 共收录 108 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 7 篇

2510.10633 2025-10-14 cs.AI 88%

Collaborative Text-to-Image Generation via Multi-Agent Reinforcement Learning and Semantic Fusion

Jiabao Shi, Minfeng Qi, Lefeng Zhang, Di Wang, Yingjie Zhao, Ziying Li, Yalong Xing, Ningran Li

机构 * Minzu University of China(民族大学) City University of Macau(澳门城市大学) Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(计算能力网络与信息安全重点实验室,教育部,山东计算机科学中心(济南国家超级计算机中心),齐鲁工业大学(山东省科学院)) The University of Adelaide(阿德莱德大学)

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract)

Comments 16 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08114 2025-10-14 cs.CV 84%

RATLIP: Generative Adversarial CLIP Text-to-Image Synthesis Based on Recurrent Affine Transformations

Chengde Lin, Xijun Lu, Guangxi Chen

机构 * School of Artificial Intelligence, Guangxi Colleges and Universities Key Laboratory of AI Algorithm Engineering(人工智能学院、广西 Colleges and Universities AI 算法工程重点实验室)

专题命中 文生图 :text-to-image(title);image synthesis(title);分类 cs.CV

Comments Accepted by 2024 IEEE International Conference on Systems, Man, and Cybernetics(SMC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01720 2025-10-14 cs.CV cs.GR cs.LG 81%

Generating Multi-Image Synthetic Data for Text-to-Image Customization

Nupur Kumari, Xi Yin, Jun-Yan Zhu, Ishan Misra, Samaneh Azadi

机构 * Carnegie Mellon University(卡内基梅隆大学) Meta

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV、cs.GR

Comments ICCV 2025. Project webpage: https://www.cs.cmu.edu/~syncd-project/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10715 2025-10-14 cs.GR cs.CV 79%

VLM-Guided Adaptive Negative Prompting for Creative Generation

Shelly Golan, Yotam Nitzan, Zongze Wu, Or Patashnik

机构 * Adobe Research(Adobe研究院) Tel Aviv University(特拉维夫大学)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV、cs.GR

Comments Project page at: https://shelley-golan.github.io/VLM-Guided-Creative-Generation/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11314 2025-10-14 cs.CL 78%

Template-Based Text-to-Image Alignment for Language Accessibility: A Study on Visualizing Text Simplifications

Belkiss Souayed, Sarah Ebling, Yingqiang Gao

机构 * University of Zurich(苏黎世大学)

专题命中 文生图 :text-to-image(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06701 2025-10-14 cs.CV cs.AI cs.LG 74%

Camouflaged Image Synthesis Is All You Need to Boost Camouflaged Detection

Haichao Zhang, Can Qin, Yu Yin, Yun Fu

机构 * Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学) Department of Electrical Engineering and Computer Science, Case Western Reserve University(电气工程与计算机科学系,凯斯西储大学)

专题命中 文生图 :image synthesis(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05970 2025-10-14 cs.CV 57%

Automatic Synthesis of High-Quality Triplet Data for Composed Image Retrieval

Haiwen Li, Delong Liu, Zhaohui Hou, Zhicheng Zhao, Fei Su

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments This paper was originally submitted to ACM MM 2025 on April 12, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 图像编辑 4 篇

2405.00313 2025-10-14 cs.CV 88%

Streamlining Image Editing with Layered Diffusion Brushes

Peyman Gholami, Robert Xiao

机构 * University of British Columbia(不列颠哥伦比亚大学)

专题命中 图像编辑 :diffusion(title,abstract);image editing(title,abstract);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:2306.00219

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06044 2025-10-14 cs.CV 85%

NEP: Autoregressive Image Editing via Next Editing Token Prediction

Huimin Wu, Xiaojian Ma, Haozhe Zhao, Yanpeng Zhao, Qing Li

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(一般人工智能国家重点实验室,BIGAI) Peking University(北京大学)

专题命中 图像编辑 :image editing(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV

Comments The project page is: https://nep-bigai.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11335 2025-10-14 cs.LG 78%

DiffStyleTS: Diffusion Model for Style Transfer in Time Series

Mayank Nagda, Phil Ostheimer, Justus Arweiler, Indra Jungjohann, Jennifer Werner, Dennis Wagner, Aparna Muraleedharan, Pouya Jafari, Jochen Schmid, Fabian Jirasek, Jakob Burger, Michael Bortz, Hans Hasse, Stephan Mandt, Marius Kloft, Sophie Fellenz

机构 * RPTU Kaiserslautern-Landau(凯斯莱特大学) Fraunhofer ITWM(弗劳恩霍夫研究所) TU Munich(慕尼黑技术大学) University of California(加州大学)

专题命中 图像编辑 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11050 2025-10-14 cs.CV 70%

Zero-shot Face Editing via ID-Attribute Decoupled Inversion

Yang Hou, Minggu Wang, Jianjun Zhao

机构 * Graduate School and Faculty of Information Science and Electrical Engineering, Kyushu University(九州大学研究生院和信息科学与电气工程学系)

专题命中 图像编辑 :diffusion(abstract);image editing(abstract);分类 cs.CV

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 扩散模型 78 篇

2510.11117 2025-10-14 cs.CV 85%

Demystifying Numerosity in Diffusion Models -- Limitations and Remedies

Yaqi Zhao, Xiaochen Wang, Li Dong, Wentao Zhang, Yuhui Yuan

机构 * Peking University(北京大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11690 2025-10-14 cs.CV cs.LG 83%

Diffusion Transformers with Representation Autoencoders

Boyang Zheng, Nanye Ma, Shengbang Tong, Saining Xie

机构 * New York University(纽约大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Technical Report; Project Page: https://rae-dit.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10918 2025-10-14 cs.CV cs.AI cs.LG 83%

DreamMakeup: Face Makeup Customization using Latent Diffusion Models

Geon Yeong Park, Inhwa Han, Serin Yang, Yeobin Hong, Seongmin Jeong, Heechan Jeon, Myeongjin Goh, Sung Won Yi, Jin Nam, Jong Chul Ye

机构 * KAIST(韩国科学技术院) Amorepacific(阿摩尔太平洋)

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07392 2025-10-14 cs.CV 83%

ID-Booth: Identity-consistent Face Generation with Diffusion Models

Darian Tomašević, Fadi Boutros, Chenhao Lin, Naser Damer, Vitomir Štruc, Peter Peer

机构 * University of Ljubljana, Faculty of Computer and Information Science(卢布尔雅纳大学计算机与信息科学学院) Fraunhofer Institute for Computer Graphics Research IGD(弗劳恩霍夫计算机图形研究机构) Xi’an Jiaotong University, School of Cyber Science and Engineering(西安交通大学网络科学与工程学院) Department of Computer Science, TU Darmstadt(图腾大学计算机科学系) University of Ljubljana, Faculty of Electrical Engineering(卢布尔雅纳大学电子工程学院)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments IEEE International Conference on Automatic Face and Gesture Recognition (FG) 2025, 14 pages

Journal ref IEEE International Conference on Automatic Face and Gesture Recognition (FG), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08632 2025-10-14 cs.CV 83%

RoboSwap: A GAN-driven Video Diffusion Framework For Unsupervised Robot Arm Swapping

Yang Bai, Liudi Yang, George Eskandar, Fengyi Shen, Dong Chen, Mohammad Altillawi, Ziyuan Liu, Gitta Kutyniok

机构 * Ludwig-Maximilians-Universität München (LMU Munich)(慕尼黑路德维希-马克西米利安大学) Huawei Heisenberg Research Center(华为海森堡研究中心) Technische Universität München (TUM)(慕尼黑技术大学) Albert-Ludwigs-Universität Freiburg(弗赖堡阿尔伯特-路易斯-大学)

专题命中 扩散模型 :diffusion(title,abstract);image editing(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12677 2025-10-14 cs.CV 83%

CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion Models

Shristi Das Biswas, Arani Roy, Kaushik Roy

机构 * Purdue University(普渡大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10149 2025-10-14 cs.LG 82%

Robust Learning of Diffusion Models with Extremely Noisy Conditions

Xin Chen, Gillian Dobbie, Xinyu Wang, Feng Liu, Di Wang, Jingfeng Zhang

机构 * University of Auckland(奥克兰大学) Shandong University(山东大学) University of Melbourne(墨尔本大学) King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14384 2025-10-14 cs.CV cs.GR 82%

Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction

Yuanhao Cai, He Zhang, Kai Zhang, Yixun Liang, Mengwei Ren, Fujun Luan, Qing Liu, Soo Ye Kim, Jianming Zhang, Zhifei Zhang, Yuqian Zhou, Yulun Zhang, Xiaokang Yang, Zhe Lin, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学) Adobe Research(Adobe研究) HKUST(香港科技大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV、cs.GR

Comments ICCV 2025; A novel one-stage 3DGS-based diffusion for 3D object generation and scene reconstruction from a single view in ~6 seconds

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11096 2025-10-14 cs.CV 79%

CoDefend: Cross-Modal Collaborative Defense via Diffusion Purification and Prompt Optimization

Fengling Zhu, Boshi Liu, Jingyu Hua, Sheng Zhong

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10910 2025-10-14 cs.CV eess.IV 79%

SceneTextStylizer: A Training-Free Scene Text Style Transfer Framework with Diffusion Model

Honghui Yuan, Keiji Yanai

机构 * The University of Electro-Communications(电通大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10868 2025-10-14 cs.CV 79%

FastHMR: Accelerating Human Mesh Recovery via Token and Layer Merging with Diffusion Decoding

Soroush Mehraban, Andrea Iaboni, Babak Taati

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所) KITE Research Institute(KITE研究 institute)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Project page: https://soroushmehraban.github.io/FastHMR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10670 2025-10-14 cs.CV 79%

AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes

Yu Li, Menghan Xia, Gongye Liu, Jianhong Bai, Xintao Wang, Conglang Zhang, Yuxuan Lin, Ruihang Chu, Pengfei Wan, Yujiu Yang

机构 * Tsinghua University(清华大学) HUST(华中科技大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) HKUST(香港科技大学) Zhejiang University(浙江大学) Wuhan University(武汉大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10663 2025-10-14 cs.CV cs.AI 79%

Scalable Face Security Vision Foundation Model for Deepfake, Diffusion, and Spoofing Detection

Gaojian Wang, Feng Lin, Tong Wu, Zhisheng Yan, Kui Ren

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 18 pages, 9 figures, project page: https://fsfm-3c.github.io/fsvfm.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10577 2025-10-14 cs.CV 79%

Injecting Frame-Event Complementary Fusion into Diffusion for Optical Flow in Challenging Scenes

Haonan Wang, Hanyu Zhou, Haoyue Liu, Luxin Yan

机构 * National Key Lab of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(多谱段信息智能处理国家级实验室,人工智能与自动化学院,华中科技大学) School of Computing, National University of Singapore(计算机学院,新加坡国立大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10434 2025-10-14 cs.CV cs.RO 79%

MonoSE(3)-Diffusion: A Monocular SE(3) Diffusion Framework for Robust Camera-to-Robot Pose Estimation

Kangjian Zhu, Haobo Jiang, Yigong Zhang, Jianjun Qian, Jian Yang, Jin Xie

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) ANGEL CorpLab and College of Computing and Data Science, Nanyang Technological University(南洋理工大学ANGEL CorpLab和计算与数据科学学院) College of Computer Science, Nankai University(南开大学计算机科学学院) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24695 2025-10-14 cs.CV cs.AI 79%

SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

Junsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu, Junyu Chen, Shuai Yang, Xianbang Wang, Yicheng Pan, Daquan Zhou, Huan Ling, Haozhe Liu, Hongwei Yi, Hao Zhang, Muyang Li, Yukang Chen, Han Cai, Sanja Fidler, Ping Luo, Song Han, Enze Xie

机构 * NVIDIA HKU(香港大学) MIT(麻省理工学院) THU(清华大学) PKU(北京大学) KAUST(国王 Abdullah 基础研究科学研究院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 21 pages, 15 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24469 2025-10-14 cs.CV cs.AI 79%

LaMoGen: Laban Movement-Guided Diffusion for Text-to-Motion Generation

Heechang Kim, Gwanghyun Kim, Se Young Chun

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08980 2025-10-14 cs.LG cs.AI cs.CV 79%

Learning Diffusion Models with Flexible Representation Guidance

Chenyu Wang, Cai Zhou, Sharut Gupta, Zongyu Lin, Stefanie Jegelka, Stephen Bates, Tommi Jaakkola

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments NeurIPS 2025; Also Oral at ICML 2025 FM4LS workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09934 2025-10-14 cs.CV cs.AI 79%

Denoising Diffusion as a New Framework for Underwater Images

Nilesh Jain, Elie Alhajjar

机构 * University of Witwatersrand(沃斯兰大学) RAND Corporation(RAND公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏