arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-10-28 至 2025-10-28 共收录 117 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 10 篇

2411.16503 2025-10-28 cs.CV 90%

Noise Diffusion for Enhancing Semantic Faithfulness in Text-to-Image Synthesis

Boming Miao, Chunxiao Li, Xiaoxiao Wang, Andi Zhang, Rui Sun, Zizhe Wang, Yao Zhu

机构 * Beijing Normal University(北京师范大学) University of Chinese Academy of Sciences(中国科学院大学) University of Manchester(曼彻斯特大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Tsinghua University(清华大学)

专题命中 文生图 :diffusion(title,abstract);text-to-image(title);image synthesis(title);分类 cs.CV

Comments Updated author formatting; no substantive changes

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23574 2025-10-28 cs.CV 89%

More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models

Hongkai Lin, Dingkang Liang, Mingyang Du, Xin Zhou, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 文生图 :text-to-image(title,abstract);diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025. The code will be made available at https://github.com/H-EmbodVis/MERGE

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22194 2025-10-28 cs.CV cs.LG 88%

ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image Generation

Yunhong Min, Daehyeon Choi, Kyeongmin Yeo, Jihyun Lee, Minhyuk Sung

机构 * KAIST(韩国科学技术院)

专题命中 文生图 :image generation(title,abstract);text-to-image(title,abstract);分类 cs.CV

Comments Project Page: https://origen2025.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21821 2025-10-28 cs.CV cs.AI 79%

Prompt fidelity of ChatGPT4o / Dall-E3 text-to-image visualisations

Dirk HR Spennemann

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22970 2025-10-28 cs.CV 70%

VALA: Learning Latent Anchors for Training-Free and Temporally Consistent

Zhangkai Wu, Xuhui Fan, Zhongyuan Xie, Kaize Shi, Longbing Cao

专题命中 文生图 :text-to-image(abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22004 2025-10-28 cs.CV 70%

LiteDiff

Ruchir Namjoshi, Nagasai Thadishetty, Vignesh Kumar, Hemanth Venkateshwara

专题命中 文生图 :diffusion(abstract);image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10776 2025-10-28 cs.CR cs.AI 67%

ME: Trigger Element Combination Backdoor Attack on Copyright Infringement

Feiyu Yang, Siyuan Liang, Aishan Liu, Dacheng Tao

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

Comments Finding unfinished issue in this work , still refining

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07555 2025-10-28 cs.CV cs.AI 57%

Synthesize Privacy-Preserving High-Resolution Images via Private Textual Intermediaries

Haoxiang Wang, Zinan Lin, Da Yu, Huishuai Zhang

机构 * Peking University(北京大学) Microsoft Research(微软研究院) Google Research(谷歌研究院) Wangxuan Institute of Computer Technology, Peking University(计算机技术研究院,北京大学) State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09744 2025-10-28 cs.CV 57%

RealCustom++: Representing Images as Real Textual Word for Real-Time Customization

Zhendong Mao, Mengqi Huang, Fei Ding, Mingcong Liu, Qian He, Yongdong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) ByteDance Inc(字节跳动公司)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments 18 pages

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08010 2025-10-28 cs.CV cs.AI 57%

Vision Transformers Don't Need Trained Registers

Nick Jiang, Amil Dravid, Alexei Efros, Yossi Gandelsman

机构 * UC Berkeley(伯克利大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV

Comments Project page and code: https://avdravid.github.io/test-time-registers. Accepted to NeurIPS '25 (spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 图像编辑 4 篇

2510.22337 2025-10-28 cs.CV 83%

GeoDiffusion: A Training-Free Framework for Accurate 3D Geometric Conditioning in Image Generation

Phillip Mueller, Talip Uenlue, Sebastian Schmidt, Marcel Kollovieh, Jiajie Fan, Stephan Guennemann, Lars Mikelsons

机构 * University of Augsburg(奥格斯堡大学) BMW Group(宝马集团) Technical University of Munich(慕尼黑技术大学) Leiden University(莱顿大学)

专题命中 图像编辑 :image generation(title,abstract);image editing(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06937 2025-10-28 cs.CV cs.AI 83%

CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing

Weiyan Xie, Han Gao, Didan Deng, Kaican Li, April Hua Liu, Yongxiang Huang, Nevin L. Zhang

专题命中 图像编辑 :image editing(title,abstract);text-to-image(abstract);分类 cs.CV

Comments Project Page: vaynexie.github.io/CannyEdit/; MindSpore Code: github.com/mindspore-lab/mindone/tree/master/examples/canny_edit; PyTorch Code: github.com/vaynexie/CannyEdit

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25776 2025-10-28 cs.CV cs.AI 77%

Editable Noise Map Inversion: Encoding Target-image into Noise For High-Fidelity Image Manipulation

Mingyu Kang, Yong Suk Choi

机构 * Department of Artificial Intelligence, University of Hanyang, Seoul, Korea(人工智能系,翰阳大学,韩国首尔) Department of Computer Science, University of Hanyang, Seoul, Korea(计算机科学系,翰阳大学,韩国首尔)

专题命中 图像编辑 :text-to-image(abstract);diffusion(abstract);image editing(abstract);分类 cs.CV

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14350 2025-10-28 cs.CV cs.AI cs.CL 70%

VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation

Shoubin Yu, Difan Liu, Ziqiao Ma, Yicong Hong, Yang Zhou, Hao Tan, Joyce Chai, Mohit Bansal

机构 * Adobe Research(Adobe研究机构) University of Michigan(密歇根大学) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 图像编辑 :diffusion(abstract);image editing(abstract);分类 cs.CV

Comments ICCV 2025; First three authors contributed equally. Project page: https://veggie-gen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 扩散模型 79 篇

2510.08273 2025-10-28 cs.CV 88%

One Stone with Two Birds: A Null-Text-Null Frequency-Aware Diffusion Models for Text-Guided Image Inpainting

Haipeng Liu, Yang Wang, Meng Wang

专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract);分类 cs.CV

Comments 27 pages, 11 figures, to appear at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16751 2025-10-28 cs.CV 85%

Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling

Erik Riise, Mehmet Onurcan Kaya, Dim P. Papadopoulos

机构 * Technical University of Denmark(丹麦技术大学) Pioneer Center for AI(先锋人工智能中心)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22366 2025-10-28 cs.CV cs.AI 83%

T2SMark: Balancing Robustness and Diversity in Noise-as-Watermark for Diffusion Models

Jindong Yang, Han Fang, Weiming Zhang, Nenghai Yu, Kejiang Chen

机构 * University of Science and Technology of China(中国科学技术大学) Anhui Province Key Laboratory of Digital Security(安徽省数字安全重点实验室) National University of Singapore(新加坡国立大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17131 2025-10-28 cs.CV cs.AI 83%

GOOD: Training-Free Guided Diffusion Sampling for Out-of-Distribution Detection

Xin Gao, Jiyao Liu, Guanghao Li, Yueming Lyu, Jianxiong Gao, Weichen Yu, Ningsheng Xu, Liang Wang, Caifeng Shan, Ziwei Liu, Chenyang Si

机构 * Nanjing University(南京大学) Fudan University(复旦大学) Carnegie Mellon University(卡内基梅隆大学) Chinese Academy of Sciences(中国科学院) Nanyang Technological University(南洋理工大学)

专题命中 扩散模型 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV

Comments 28 pages, 16 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14036 2025-10-28 cs.LG cs.AI 82%

Adaptive Inference-Time Scaling via Cyclic Diffusion Search

Gyubin Lee, Truong Nhat Nguyen Bao, Jaesik Yoon, Dongwoo Lee, Minsu Kim, Yoshua Bengio, Sungjin Ahn

机构 * KAIST(韩国科学技术院) Mila – Quebec AI Institute(魁北克AI研究所) Université de Montréal(蒙特利尔大学) SAP(SAP公司)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15857 2025-10-28 cs.LG cs.AI cs.CV cs.RO 80%

Diffusion Beats Autoregressive in Data-Constrained Settings

Mihir Prabhudesai, Mengning Wu, Amir Zadeh, Katerina Fragkiadaki, Deepak Pathak

机构 * Carnegie Mellon University(卡内基梅隆大学) Lambda

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Project Webpage: https://diffusion-scaling.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23382 2025-10-28 cs.CV 79%

An Efficient Remote Sensing Super Resolution Method Exploring Diffusion Priors and Multi-Modal Constraints for Crop Type Mapping

Songxi Yang, Tang Sui, Qunying Huang

机构 * Department of Geography(地理系) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 41 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21264 2025-10-28 cs.CV 79%

Topology Sculptor, Shape Refiner: Discrete Diffusion Model for High-Fidelity 3D Meshes Generation

Kaiyu Song, Hanjiang Lai, Yaqing Zhang, Chuangjian Cai, Yan Pan Kun Yue, Jian Yin

机构 * Sun Yat-sen University(中山大学) Tencent VisVise(腾讯VisVise) Yunnan University(云南大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22605 2025-10-28 cs.CV physics.med-ph 79%

Projection Embedded Diffusion Bridge for CT Reconstruction from Incomplete Data

Yuang Wang, Pengfei Jin, Siyeop Yoon, Matthew Tivnan, Shaoyang Zhang, Li Zhang, Quanzheng Li, Zhiqiang Chen, Dufan Wu

机构 * The Department of Engineering Physics, Tsinghua University(清华大学工程物理系) Center for Advanced Medical Computing and Analysis, Massachusetts General Hospital and Harvard Medical School(麻省总医院和哈佛医学院高级医学计算与分析中心) Department of Radiology, The Ohio State University(俄亥俄州立大学放射科)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 53 pages, 7 figures, submitted to Medical Image Analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22236 2025-10-28 cs.CV 79%

DiffusionLane: Diffusion Model for Lane Detection

Kunyang Zhou, Yeqin Shao

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22229 2025-10-28 cs.CV 79%

Diffusion-Driven Two-Stage Active Learning for Low-Budget Semantic Segmentation

Jeongin Kim, Wonho Bae, YouLee Han, Giyeong Oh, Youngjae Yu, Danica J. Sutherland, Junhyug Noh

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20819 2025-10-28 cs.CV cs.AI cs.LG 79%

Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge

Nimrod Berman, Omkar Joglekar, Eitan Kosman, Dotan Di Castro, Omri Azencot

机构 * Bosch AI Center(博世人工智能中心) Ben-Gurion University of the Negev(巴伊兰大学) Technical University of Munich(慕尼黑技术大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepted as a poster at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00617 2025-10-28 eess.IV cs.CV 79%

Continuous and complete liver vessel segmentation with graph-attention guided diffusion

Xiaotong Zhang, Alexander Broersen, Gonnie CM van Erp, Silvia L. Pintea, Jouke Dijkstra

机构 * Radiology department, Leiden University Medical Center(莱顿大学医学中心放射科)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepted by Knowledge-Based Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22166 2025-10-28 eess.IV cs.CV cs.LG 79%

Expert Validation of Synthetic Cervical Spine Radiographs Generated with a Denoising Diffusion Probabilistic Model

Austin A. Barr, Brij S. Karmur, Anthony J. Winder, Eddie Guo, John T. Lysack, James N. Scott, William F. Morrish, Muneer Eesa, Morgan Willson, David W. Cadotte, Michael M. H. Yang, Ian Y. M. Chan, Sanju Lama, Garnette R. Sutherland

机构 * Cumming School of Medicine, University of Calgary(卡里尔大学医学学院) Division of Neurosurgery, Department of Clinical Neurosciences, University of Calgary(卡里尔大学临床神经科学系神经外科部) Department of Radiology, University of Calgary(卡里尔大学放射科) Division of Neurosurgery, Department of Surgery, University of Toronto(多伦多大学外科系神经外科部) Department of Medical Imaging, University of Toronto(多伦多大学医学影像科) Project neuroArm, Department of Clinical Sciences, University of Calgary(卡里尔大学临床科学部neuroArm项目)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 10 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21802 2025-10-28 cs.CV cs.LG stat.ML 79%

It Takes Two to Tango: Two Parallel Samplers Improve Quality in Diffusion Models for Limited Steps

Pedro Cisneros-Velarde

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21791 2025-10-28 cs.CV physics.ins-det 79%

Exploring the design space of diffusion and flow models for data fusion

Niraj Chaudhari, Manmeet Singh, Naveen Sudharsan, Amit Kumar Srivastava, Harsh Kamath, Dushyant Mahajan, Ayan Paul

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏