arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-10-28 至 2025-10-28 共收录 117 信号源:cs.CV, cs.GR, cs.MM

1. 扩散模型 79 篇

2510.21776 2025-10-28 physics.chem-ph cond-mat.soft 50%

Tuning laser-induced optical breakdown and cavitation through the ionic environment in aqueous media

Junhao Cai, Yuhan Li, Yunqiao Liu, Benlong Wang, Mingbo Li

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14806 2025-10-28 q-bio.NC cs.LG stat.ML 50%

Place Cells as Multi-Scale Position Embeddings: Random Walk Transition Kernels for Path Planning

Minglu Zhao, Dehong Xu, Deqian Kong, Wen-Hao Zhang, Ying Nian Wu

机构 * UCLA(加州大学洛杉矶分校) UT Southwestern Medical Center(德克萨斯大学西南医学中心)

专题命中 扩散模型 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10521 2025-10-28 physics.optics quant-ph 50%

Reconfigurable quantum photonic circuits based on quantum dots

Adam McCaw, Jacob Ewaniuk, Bhavin J. Shastri, Nir Rotenberg

专题命中 扩散模型 :diffusion(abstract)

Comments manuscript includes 10 pages, 4 figures followed by the supplement with 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 可控生成 9 篇

2510.22851 2025-10-28 cs.CV cs.AI 85%

Semantic Surgery: Zero-Shot Concept Erasure in Diffusion Models

Lexiang Xiong, Chengyu Liu, Jingwen Ye, Yan Liu, Yuecong Xu

机构 * National University of Singapore(国立新加坡大学) Sichuan University(四川大学)

专题命中 可控生成 :diffusion(title,abstract);image generation(abstract);text-to-image(abstract);分类 cs.CV

Comments Accepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025). Code is available at https://github.com/Lexiang-Xiong/Semantic-Surgery

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21763 2025-10-28 cs.CV cs.AI 85%

Proportion and Perspective Control for Flow-Based Image Generation

Julien Boudier, Hugo Caselles-Dupré

机构 * Obvious Research(Obvious研究)

专题命中 可控生成 :image generation(title);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)

Comments Technical report after open-source release

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08441 2025-10-28 cs.CV cs.AI 74%

Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

Anlin Zheng, Xin Wen, Xuanyang Zhang, Chuofan Ma, Tiancai Wang, Gang Yu, Xiangyu Zhang, Xiaojuan Qi

机构 * The University of Hong Kong(香港大学) StepFun Dexmal MEGVII Technology(MEGVII科技)

专题命中 可控生成 :image generation(title);分类 cs.CV

Comments 20 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22010 2025-10-28 cs.CV cs.LG eess.IV 70%

FlowOpt: Fast Optimization Through Whole Flow Processes for Training-Free Editing

Or Ronai, Vladimir Kulikov, Tomer Michaeli

专题命中 可控生成 :diffusion(abstract);image editing(abstract);分类 cs.CV

Comments Project's webpage at https://orronai.github.io/FlowOpt/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10999 2025-10-28 cs.CV cs.AI cs.CL cs.MM 62%

ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotations

Bowen Jiang, Yuan Yuan, Xinyi Bai, Zhuoqun Hao, Alyson Yin, Yaojie Hu, Wenyu Liao, Lyle Ungar, Camillo J. Taylor

机构 * University of Pennsylvania(宾夕法尼亚大学) Cornell University(康奈尔大学) University of California, Irvine(加州大学伊藤分校)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV、cs.MM

Comments The 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP) Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22981 2025-10-28 cs.AI cs.CV 57%

Exploring Semantic-constrained Adversarial Example with Instruction Uncertainty Reduction

Jin Hu, Jiakai Wang, Linna Jing, Haolin Li, Haodong Liu, Haotong Qin, Aishan Liu, Ke Xu, Xianglong Liu

机构 * State Key Laboratory of Complex & Critical Software Environment (CCSE), Beihang University(复杂与关键软件环境国家重点实验室,北京航空航天大学) Zhongguancun Laboratory(中关村实验室) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) Department of Information Technology and Electrical Engineering, ETH Zurich(苏黎世联邦理工学院信息科技与电气工程系)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22534 2025-10-28 cs.CV 57%

SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning

Chen Chen, Majid Abdolshah, Violetta Shevchenko, Hongdong Li, Chang Xu, Pulak Purkait

机构 * Amazon(亚马逊公司) The University of Sydney(悉尼大学) Australian National University(澳大利亚国立大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21887 2025-10-28 cs.CV cs.AI cs.LG 57%

Generative AI in Depth: A Survey of Recent Advances, Model Variants, and Real-World Applications

Shamim Yazdani, Akansha Singh, Nripsuta Saxena, Zichong Wang, Avash Palikhe, Deng Pan, Umapada Pal, Jie Yang, Wenbin Zhang

机构 * Florida International University(佛罗里达国际大学) Manipal Institute of Technology(曼海姆技术学院) University of Southern California(南加州大学) Indian Statistical Institute(印度统计研究所) University of Wollongong(沃林根大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

Comments Accepted by the Journal of Big Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12650 2025-10-28 cs.LG cs.AI 50%

Score-informed Neural Operator for Enhancing Ordering-based Causal Discovery

Jiyeon Kang, Songseong Kim, Chanhui Lee, Doyeong Hwang, Joanie Hayoun Chung, Yunkyung Ko, Sumin Lee, Sungwoong Kim, Sungbin Lim

专题命中 可控生成 :diffusion(abstract)

Comments Accepted to NeurIPS 2025. 36 pages, 18 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 图像修复 2 篇

2507.05604 2025-10-28 cs.CV eess.IV 70%

Kernel Density Steering: Inference-Time Scaling via Mode Seeking for Image Restoration

Yuyang Hu, Kangfu Mei, Mojtaba Sahraee-Ardakan, Ulugbek S. Kamilov, Peyman Milanfar, Mauricio Delbracio

机构 * Google(谷歌) Washington University in St. Louis(圣路易斯华盛顿大学)

专题命中 图像修复 :diffusion(abstract);inpainting(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.11672 2025-10-28 cs.CV 57%

Invertible generative models for inverse problems: mitigating representation error and dataset bias

Muhammad Asim, Mara Daniels, Oscar Leong, Ali Ahmed, Paul Hand

机构 * itu(信息科技大学) rice(里士满大学) neumcs(东北大学)

专题命中 图像修复 :inpainting(abstract);分类 cs.CV

Comments Camera ready version for ICML 2020, paper 2655

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 个性化与一致性 6 篇

2510.22994 2025-10-28 cs.CV 70%

SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency

Quanjian Song, Donghao Zhou, Jingyu Lin, Fei Shen, Jiaze Wang, Xiaowei Hu, Cunjian Chen, Pheng-Ann Heng

机构 * Monash University(莫纳什大学) The Chinese University of Hong Kong(香港中文大学) National University of Singapore(新加坡国立大学) South China University of Technology(华南理工大学)

专题命中 个性化与一致性 :image generation(abstract);text-to-image(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025; Project Page: https://lulupig12138.github.io/SceneDecorator

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23605 2025-10-28 cs.CV cs.AI cs.GR cs.LG cs.RO 62%

Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling

Shuhong Zheng, Ashkan Mirzaei, Igor Gilitschenski

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所) Snap Inc.(Snap公司)

专题命中 个性化与一致性 :inpainting(abstract);分类 cs.CV、cs.GR

Comments NeurIPS 2025, 38 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22810 2025-10-28 cs.CV 57%

MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control

Fatemeh Nazarieh, Zhenhua Feng, Diptesh Kanojia, Muhammad Awais, Josef Kittler

机构 * School of Computer Science and Electronic Engineering, University of Surrey(Surrey大学计算机科学与电子工程学院) School of Artificial Intelligence and Computer Science, Jiangnan University(江南大学人工智能与计算机科学学院) Centre for Vision, Speech and Signal Processing, University of Surrey(Surrey大学视觉、语音和信号处理中心)

专题命中 个性化与一致性 :diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04705 2025-10-28 cs.CV 57%

Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations

Yuji Wang, Moran Li, Xiaobin Hu, Ran Yi, Jiangning Zhang, Han Feng, Weijian Cao, Yabiao Wang, Chengjie Wang, Lizhuang Ma

机构 * Shanghai Jiao Tong University, Tencent Youtu Lab(上海交通大学,腾讯云图实验室) Tencent Youtu Lab(腾讯云图实验室) Shanghai Jiao Tong University(上海交通大学) Tencent(腾讯)

专题命中 个性化与一致性 :text-to-image(abstract);分类 cs.CV

Comments ACM Multimedia 2025; code URL: https://github.com/rain152/IPVG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21857 2025-10-28 cs.CV cs.AI 57%

Poisson Flow Consistency Training

Anthony Zhang, Mahmut Gokmen, Dennis Hein, Rongjun Ge, Wenjun Xia, Ge Wang, Jin Chen

机构 * Pratt School of Engineering Duke University(普拉特工程学院 哥伦比亚大学) Biomedical Imaging Center Rensselaer Polytechnic Institute(生物医学成像中心 莱文森理工学院) Department of Computer Science University of Kentucky(计算机科学系 肯塔基大学) Department of Medicine Department of Biomedical Informations University of Alabama at Birmingham(医学系 生物医学信息学系 亚拉巴马大学伯明翰分校) School of Engineering Sciences Department of Physics KTH Royal Institute of Technology(工程科学学院 物理系 瑞典皇家理工学院)

专题命中 个性化与一致性 :image generation(abstract);分类 cs.CV

Comments 5 pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23106 2025-10-28 cs.LG 50%

Sampling from Energy distributions with Target Concrete Score Identity

Sergei Kholkin, Francisco Vargas, Alexander Korotin

机构 * Applied AI Institute(应用人工智能研究所)

专题命中 个性化与一致性 :diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 图像生成评测 4 篇

2510.23020 2025-10-28 cs.CV cs.CL 85%

M$^{3}$T2IBench: A Large-Scale Multi-Category, Multi-Instance, Multi-Relation Text-to-Image Benchmark

Huixuan Zhang, Xiaojun Wan

机构 * Wangxuan Institute of Computer Technology, Peking University(计算机技术研究所,北京大学)

专题命中 图像生成评测 :text-to-image(title,abstract);image generation(abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22521 2025-10-28 cs.CV cs.AI cs.IR cs.LG 79%

Open Multimodal Retrieval-Augmented Factual Image Generation

Yang Tian, Fan Liu, Jingyuan Zhang, Wei Bi, Yupeng Hu, Liqiang Nie

机构 * Shandong University(山东大学) National University of Singapore(新加坡国立大学) Kuaishou Technology(快手科技) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 图像生成评测 :image generation(title,abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21347 2025-10-28 cs.LG 78%

OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models

Ziheng Cheng, Yixiao Huang, Hui Xu, Somayeh Sojoudi, Xuandong Zhao, Dawn Song, Song Mei

机构 * UC Berkeley(加州大学伯克利分校)

专题命中 图像生成评测 :text-to-image(title,abstract)

Comments NeurIPS 2025 (D&B Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23023 2025-10-28 cs.CV cs.CL 77%

UniAIDet: A Unified and Universal Benchmark for AI-Generated Image Content Detection and Localization

Huixuan Zhang, Xiaojun Wan

机构 * Wangxuan Institute of Computer Technology(计算机技术研究所)

专题命中 图像生成评测 :text-to-image(abstract);image editing(abstract);inpainting(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 效率与蒸馏 2 篇

2503.10270 2025-10-28 cs.CV 79%

EEdit: Rethinking the Spatial and Temporal Redundancy for Efficient Image Editing

Zexuan Yan, Yue Ma, Chang Zou, Wenteng Chen, Qifeng Chen, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Hong Kong University of Science and Technology(香港科技大学)

专题命中 效率与蒸馏 :image editing(title,abstract);分类 cs.CV

Comments accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01320 2025-10-28 cs.LG cs.AI cs.CV 57%

Psi-Sampler: Initial Particle Sampling for SMC-Based Inference-Time Reward Alignment in Score Models

Taehoon Yoon, Yunhong Min, Kyeongmin Yeo, Minhyuk Sung

机构 * KAIST(韩国科学技术院)

专题命中 效率与蒸馏 :image generation(abstract);分类 cs.CV

Comments NeurIPS 2025, Spotlight Presentation

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他图像生成 1 篇

2510.21815 2025-10-28 eess.IV cs.AI cs.CV 57%

HDR Image Reconstruction using an Unsupervised Fusion Model

Kumbha Nagaswetha

机构 * Department of Electrical Communication Engineering, Indian Institute of Science(电子通信工程系,印度科学研究院)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏