arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-09-22 至 2025-09-22 共收录 58 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 7 篇

2509.15962 2025-09-22 cs.AI 87%

Structured Information for Improving Spatial Relationships in Text-to-Image Generation

Sander Schildermans, Chang Tian, Ying Jiao, Marie-Francine Moens

专题命中 文生图 :text-to-image(title,abstract);image generation(title,comments)

Comments text-to-image generation, structured information, spatial relationship

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16141 2025-09-22 cs.CV 83%

AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models

Vatsal Malaviya, Agneet Chatterjee, Maitreya Patel, Yezhou Yang, Chitta Baral

机构 * Arizona State University(亚利桑那州立大学)

专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);分类 cs.CV

Comments Project Page : https://vatsal-malaviya.github.io/AcT2I/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15803 2025-09-22 cs.CV cs.AI 79%

CIDER: A Causal Cure for Brand-Obsessed Text-to-Image Models

Fangjian Shen, Zifeng Liang, Chao Wang, Wushao Wen

机构 * School of Computer Science(计算机科学学院) Engineering, Sun Yat-sen University, Guangzhou, China(工程学院,中山大学,广州,中国)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments 5 pages, 7 figures, submitted to ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16197 2025-09-22 cs.CV cs.CL cs.LG 77%

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

Yanghao Li, Rui Qian, Bowen Pan, Haotian Zhang, Haoshuo Huang, Bowen Zhang, Jialing Tong, Haoxuan You, Xianzhi Du, Zhe Gan, Hyunjik Kim, Chao Jia, Zhenbang Wang, Yinfei Yang, Mingfei Gao, Zi-Yi Dou, Wenze Hu, Chang Gao, Dongxu Li, Philipp Dufter, Zirui Wang, Guoli Yin, Zhengdong Zhang, Chen Chen, Yang Zhao, Ruoming Pang, Zhifeng Chen

机构 * Apple(苹果公司)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15270 2025-09-22 cs.CV cs.AI 70%

PRISM: Phase-enhanced Radial-based Image Signature Mapping framework for fingerprinting AI-generated images

Emanuele Ricco, Elia Onofri, Lorenzo Cima, Stefano Cresci, Roberto Di Pietro

专题命中 文生图 :text-to-image(abstract);diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15391 2025-09-22 cs.CV 57%

RaceGAN: A Framework for Preserving Individuality while Converting Racial Information for Image-to-Image Translation

Mst Tasnim Pervin, George Bebis, Fang Jiang, Alireza Tavakkoli

机构 * 1 2 4 Department of Computer Science \& Engineering, 3 Department of Psychology, University of Nevada, Reno, USA Email: 1 , 2 , 3 , 4

专题命中 文生图 :image synthesis(abstract);分类 cs.CV

Journal ref ICMLA 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15432 2025-09-22 cs.IR 50%

SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models

Thong Nguyen, Yibin Lei, Jia-Huei Ju, Andrew Yates

专题命中 文生图 :text-to-image(abstract)

Comments Accepted

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 扩散模型 41 篇

2410.08567 2025-09-22 cs.CV 88%

Diffusion-Based Depth Inpainting for Transparent and Reflective Objects

Tianyu Sun, Dingchang Hu, Yixiang Dai, Guijin Wang

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 扩散模型 :diffusion(title,abstract);inpainting(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15678 2025-09-22 cs.CV 83%

Layout Stroke Imitation: A Layout Guided Handwriting Stroke Generation for Style Imitation with Diffusion Model

Sidra Hanif, Longin Jan Latecki

机构 * Temple University, Philadelphia PA, USA(特拉华大学)

专题命中 扩散模型 :diffusion(title,abstract);image generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16106 2025-09-22 eess.IV cs.CV cs.LG 79%

PRISM: Probabilistic and Robust Inverse Solver with Measurement-Conditioned Diffusion Prior for Blind Inverse Problems

Yuanyun Hu, Evan Bell, Guijin Wang, Yu Sun

机构 * Johns Hopkins University(约翰霍普金斯大学) Tsinghua University(清华大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16091 2025-09-22 cs.CV 79%

Blind-Spot Guided Diffusion for Self-supervised Real-World Denoising

Shen Cheng, Haipeng Li, Haibin Huang, Xiaohong Liu, Shuaicheng Liu

机构 * Dexmal UESTC Tele AI SJTU

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16019 2025-09-22 eess.IV cs.CV 79%

SLaM-DiMM: Shared Latent Modeling for Diffusion Based Missing Modality Synthesis in MRI

Bhavesh Sandbhor, Bheeshm Sharma, Balamurugan Palaniappan

机构 * IIT Bombay(印度理工学院班加罗尔分校)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15553 2025-09-22 cs.CV cs.AI stat.AP 79%

Diffusion-Based Cross-Modal Feature Extraction for Multi-Label Classification

Tian Lan, Yiming Zheng, Jianxin Yin

机构 * School of Statistics, Renmin University of China(中国人民大学统计学院) Center for Applied Statistics and School of Statistics, Renmin University of China(中国人民大学应用统计中心和统计学院)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13922 2025-09-22 cs.CV 79%

Towards Robust Defense against Customization via Protective Perturbation Resistant to Diffusion-based Purification

Wenkui Yang, Jie Cao, Junxian Duan, Ran He

机构 * MAIS & NLPR, Institute of Automation, Chinese Academy of Sciences(自动化研究所、中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院、中国科学院大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08437 2025-09-22 cs.CV 79%

TT-DF: A Large-Scale Diffusion-Based Dataset and Benchmark for Human Body Forgery Detection

Wenkui Yang, Zhida Zhang, Xiaoqiang Zhou, Junxian Duan, Jie Cao

机构 * MAIS \& NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing, China School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China University of Science

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepted by PRCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06793 2025-09-22 eess.IV cs.CV 79%

HistDiST: Histopathological Diffusion-based Stain Transfer

Erik Großkopf, Valay Bundele, Mehran Hosseinzadeh, Hendrik P. A. Lensch

机构 * University of Tübingen(图宾根大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments Accepted to DAGM GCPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18104 2025-09-22 cs.CV 79%

PromptMID: Modal Invariant Descriptors Based on Diffusion and Vision Foundation Models for Optical-SAR Image Matching

Han Nie, Bin Luo, Jun Liu, Zhitao Fu, Huan Zhou, Shuo Zhang, Weixing Liu

机构 * State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(信息工程测绘与遥感国家重点实验室,武汉大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

Comments 15 pages, 8 figures

Journal ref ISPRS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14710 2025-09-22 cs.CV cs.AI cs.LG 79%

G2D2: Gradient-Guided Discrete Diffusion for Inverse Problem Solving

Naoki Murata, Chieh-Hsin Lai, Yuhta Takida, Toshimitsu Uesaka, Bac Nguyen, Stefano Ermon, Yuki Mitsufuji

机构 * Sony AI(索尼人工智能) Stanford University(斯坦福大学) Sony Group Corporation(索尼集团)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15427 2025-09-22 physics.app-ph 78%

Thin-film boundary-layer diffusion of non-equilibrium flow to kinetically limited reactive surfaces via Damköhler thermochemistry tables

Jeffrey D. Engerer, Lincoln N. Collins

专题命中 扩散模型 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15415 2025-09-22 cond-mat.stat-mech 78%

Spectral Characterization of Wave Scattering at a Granular-Elastic Solid Interface: From Hyperbolic Wave Propagation to Near-Parabolic Diffusion

Joshua R. Tempelman, Chongan Wang, Alexander F. Vakakis

专题命中 扩散模型 :diffusion(title,abstract)

Comments 18 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14951 2025-09-22 math.PR 78%

Stochastic Hamiltonian Type Jump Diffusion Systems with Countable Regimes: Strong Feller Property and Exponential Ergodicity

Fubao Xi, Yafei Zhai, Zuozheng Zhang

专题命中 扩散模型 :diffusion(title,abstract)

Comments arXiv admin note: text overlap with arXiv:1702.01048

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14082 2025-09-22 cs.RO 78%

FlightDiffusion: Revolutionising Autonomous Drone Training with Diffusion Models Generating FPV Video

Valerii Serpiva, Artem Lykov, Faryal Batool, Vladislav Kozlovskiy, Miguel Altamirano Cabrera, Dzmitry Tsetserukou

机构 * Intelligent Space Robotics Laboratory, Skolkovo Institute of Science and Technology Moscow(斯克尔科夫科学与技术研究所莫斯科智能空间机器人实验室)

专题命中 扩散模型 :diffusion(title,abstract)

Comments Submitted to conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15311 2025-09-22 cs.IR 78%

Modeling Long-term User Behaviors with Diffusion-driven Multi-interest Network for CTR Prediction

Weijiang Lai, Beihong Jin, Yapeng Zhang, Yiyuan Zheng, Rui Zhao, Jian Dong, Jun Lei, Xingxing Wang

专题命中 扩散模型 :diffusion(title,abstract)

Journal ref RecSys 2025: Proceedings of the Nineteenth ACM Conference on Recommender Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17173 2025-09-22 math.PR math.AP 78%

Existence results for the Cox-Ingersoll-Ross model with variable exponent diffusion

Mustafa Avci

专题命中 扩散模型 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07318 2025-09-22 cs.SD cs.AI eess.AS 78%

Generating Moving 3D Soundscapes with Latent Diffusion Models

Christian Templin, Yanda Zhu, Hao Wang

专题命中 扩散模型 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05847 2025-09-22 cs.NE 78%

A Universal Framework for Large-Scale Multi-Objective Optimization Based on Particle Drift and Diffusion

Jia-Cheng Li, Min-Rong Chen, Guo-Qiang Zeng, Jian Weng, Man Wang, Jia-Lin Mai

专题命中 扩散模型 :diffusion(title,abstract)

Comments There are several details related to operators are imprecise.To uphold the principle of accuracy, we have decided to retract the article for now

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21943 2025-09-22 physics.bio-ph q-bio.QM 78%

Single-Trajectory Bayesian Modeling Reveals Multi-State Diffusion of the MSH Sliding Clamp

Seongyu Park, Inho Yang, Jinseob Lee, Sinwoo Kim, Juana Martín-López, Richard Fishel, Jong-Bong Lee, Jae-Hyung Jeon

专题命中 扩散模型 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13759 2025-09-22 cs.LG cs.AI 78%

Discrete Diffusion in Large Language and Multimodal Models: A Survey

Runpeng Yu, Qi Li, Xinchao Wang

专题命中 扩散模型 :diffusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09381 2025-09-22 eess.AS cs.SD 78%

DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers

Heitor R. Guimarães, Jiaqi Su, Rithesh Kumar, Tiago H. Falk, Zeyu Jin

机构 * INRS-EMT, Université du Québec, Montréal, Canada(INRS-EMT,魁北克大学,蒙特利尔,加拿大) Adobe Research, San Francisco, California, United States(Adobe研究,旧金山,加利福尼亚,美国)

专题命中 扩散模型 :diffusion(title,abstract)

Comments Manuscript under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01064 2025-09-22 cs.CV cs.AI cs.LG cs.MM eess.IV 62%

FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait

Taekyung Ki, Dongchan Min, Gyeongsu Chae

机构 * KAIST(韩国科学技术院) DeepBrain AI Inc.(DeepBrain AI公司)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV、cs.MM

Comments ICCV 2025. Project page: https://deepbrainai-research.github.io/float/

详情

展开后加载摘要…

URL PDF HTML 收藏