arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 3486 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3486 篇

2004.03590 2020-04-09 cs.CV cs.GR cs.LG cs.NE eess.IV 81%

Multimodal Image Synthesis with Conditional Implicit Maximum Likelihood Estimation

Ke Li, Shichong Peng, Tianhao Zhang, Jitendra Malik

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV、cs.GR

Comments To appear in International Journal of Computer Vision (IJCV). arXiv admin note: text overlap with arXiv:1811.12373

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.12356 2019-04-30 cs.CV cs.GR 81%

Deferred Neural Rendering: Image Synthesis using Neural Textures

Justus Thies, Michael Zollhöfer, Matthias Nießner

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV、cs.GR

Comments Video: https://youtu.be/z-pVip6WeyY SIGGRAPH 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.11585 2018-08-21 cs.CV cs.GR cs.LG 81%

High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs

Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, Bryan Catanzaro

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV、cs.GR

Comments v2: CVPR camera ready, adding more results for edge-to-photo examples

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.10992 2018-05-01 cs.CV cs.AI cs.GR cs.LG 81%

Semi-parametric Image Synthesis

Xiaojuan Qi, Qifeng Chen, Jiaya Jia, Vladlen Koltun

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV、cs.GR

Comments Published at the Conference on Computer Vision and Pattern Recognition (CVPR 2018)

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.02823 2018-04-17 cs.CV cs.GR 81%

TextureGAN: Controlling Deep Image Synthesis with Texture Patches

Wenqi Xian, Patsorn Sangkloy, Varun Agrawal, Amit Raj, Jingwan Lu, Chen Fang, Fisher Yu, James Hays

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV、cs.GR

Comments CVPR 2018 spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.05349 2017-08-18 cs.CV cs.GR cs.LG 81%

PixelNN: Example-based Image Synthesis

Aayush Bansal, Yaser Sheikh, Deva Ramanan

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV、cs.GR

Comments Project Page: http://www.cs.cmu.edu/~aayushb/pixelNN/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01888 2026-04-03 cs.CV 80%

Low-Effort Jailbreak Attacks Against Text-to-Image Safety Filters

低努力对抗攻击针对文本到图像安全过滤器

Ahmed B Mustafa, Zihan Ye, Yang Lu, Michael P Pound, Shreyank N Gowda

机构 * University of Nottingham(诺丁汉大学) Xi’an Jiaotong-Liverpool University(西交利物浦大学) Xiamen University(厦门大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 研究揭示文本到图像模型易受低努力对抗攻击影响,通过提示策略绕过安全过滤,展示多种视觉对抗技术,揭示表面提示过滤与深层语义理解的差距,攻击成功率高达74.47%。

Comments Text-to-Image version of the Anyone can Jailbreak paper. Accepted in CVPR-W AIMS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13619 2023-11-27 cs.CV cs.CR 80%

Steal My Artworks for Fine-tuning? A Watermarking Framework for Detecting Art Theft Mimicry in Text-to-Image Models

Ge Luo, Junqiang Huang, Manman Zhang, Zhenxing Qian, Sheng Li, Xinpeng Zhang

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments A Watermarking Framework for Detecting Art Theft Mimicry in Text-to-Image Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.15678 2021-12-10 cs.CV 80%

A Shading-Guided Generative Implicit Model for Shape-Accurate 3D-Aware Image Synthesis

Xingang Pan, Xudong Xu, Chen Change Loy, Christian Theobalt, Bo Dai

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS2021. We proposed ShadeGAN, which could perform shape-accurate 3D-aware image synthesis by modeling shading in generative implicit models

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.12666 2021-08-10 cs.CV 80%

Semantically Self-Aligned Network for Text-to-Image Part-aware Person Re-identification

Zefeng Ding, Changxing Ding, Zhiyin Shao, Dacheng Tao

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

Comments A new database for text-to-image ReID is provided. Code will be released

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06761 2025-05-16 cs.LG cs.MA 80%

Learning Graph Representation of Agent Diffusers

Youcef Djenouri, Nassim Belmecheri, Tomasz Michalak, Jan Dubiński, Ahmed Nabil Belbachir, Anis Yazidi

机构 * University of South-Eastern Norway(南欧挪威大学) Norwegian Research Centre(挪威研究中心) Simula Laboratory Research(Simula实验室研究) University of Warsaw(华沙大学) Warsaw University of Technology(华沙理工大学) University of Oslo(奥斯陆大学)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);diffusion(abstract);image synthesis(abstract)

Comments Accepted at AAMAS2025 International Conference on Autonomous Agents and Multiagent Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.08042 2022-01-21 cs.IR cs.LG 80%

GAN-based Matrix Factorization for Recommender Systems

Ervin Dervishaj, Paolo Cremonesi

专题命中 文生图 :image generation(abstract);text-to-image(abstract);inpainting(abstract);image synthesis(abstract)

Comments Accepted at the 37th ACM/SIGAPP Symposium on Applied Computing (SAC '22)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21784 2026-08-25 cs.CV 新提交 79%

DefaultShift: Auditing Semantic Default Shift in Accelerated Text-to-Image Models

DefaultShift:审计加速文本到图像模型中的语义默认偏移

Xuanhua Yin, Chuanzhi Xu, Shunqi Mao, Wei Guo, Weidong Cai

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 本研究提出DefaultShift审计方法检测文本到图像模型加速时的语义默认偏移,还提出DefaultShift-Select校准方法降低偏移,提升模型语义保护能力。

Comments 21 pages, 12 figures, 25 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05853 2026-08-21 cs.CV 版本更新 79%

Reading Between the Pixels: An Inscriptive Jailbreak Attack on Text-to-Image Models

在像素之间阅读:一种针对文本到图像模型的 inscription 性 jailbreak 攻击

Zonghao Ying, Haowen Dai, Lianyu Hu, Zonglei Jing, Quanchen Zou, Yaodong Yang, Aishan Liu, Xianglong Liu

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 本文提出 Etch 框架,通过分解对抗提示为语义伪装、视觉空间锚定和字形编码三层,实现对文本到图像模型的 inscription 性 jailbreak 攻击,验证了现有安全机制的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07687 2026-08-21 eess.IV cs.CV 版本更新 79%

FermatSyn: SAM2-Enhanced Bidirectional Mamba with Isotropic Spiral Scanning for Multi-Modal Medical Image Synthesis

FermatSyn: 基于改进双向Mamba的多模态医学图像合成方法

Feng Yuan, Yifan Gao, Haoyue Li, Xin Gao

机构 * USTC(中国科学技术大学) SII(上海信息研究所)

专题命中 文生图 :image synthesis(title,abstract);分类 cs.CV

AI总结 FermatSyn通过改进的双向Mamba结合Fermat螺旋扫描策略,解决多模态医学图像合成中全局一致性与局部细节的平衡问题,提升合成图像质量与临床应用价值。

Comments MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14976 2026-08-18 cs.CV 新提交 79%

Benchmarking Frontier Text-to-Image Models on Image-Description Prompts

基于图像-描述提示的前沿文本到图像模型基准测试

Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmed Rashad

机构 * Perle(珀尔)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 该研究针对组合要求高的图像-描述提示,评估了四个前沿文本到图像模型的性能,发现 Gemini 3 Pro Image 表现最优,领先系统的主要问题是对象计数错误和几何伪影。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11452 2026-08-13 cs.CV cs.AI 新提交 79%

TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation

TangPoetryBench:面向诗歌到图像生成的多维度基准与基于评分规则的评估器

Haoqi Hu, Tongji Luo, Li Zhang, Boning Zhou

专题命中 文生图 :image generation(title);text-to-image(abstract);分类 cs.CV

AI总结 该研究推出TangPoetryBench多维度基准与PAE评估器,解决T2I模型生成诗歌插图的评估难题,PAE性能接近Claude且可泛化,相关资源已公开。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03357 2026-08-05 cs.CV 新提交 79%

Can Text-to-Image Models Draw from the Right Frame of Reference?

文本到图像模型能否从正确的参考框架中生成图像?

Zheyuan Gu, Ruihang Li, Yong Huang, Yiqian Zhang, XIangzhao Hao, Jiaxin Niu, Jiahao Hu, Zhenyu Zhang

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 该研究引入FoR-T2I基准,发现现有22个T2I模型在参考框架提示下的布局理解准确率远低于相机视图提示,还提出VLM门控重写方法提升了参考框架下的生成准确率。

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00598 2026-08-04 cs.MM cs.CY 新提交 79%

EmergencyBias: Bias in Text-to-Image Models under Emergency Scenarios

EmergencyBias:文本到图像模型在应急场景下的偏差

Haibo Tang, Linqi Zhang, Hongxin Huan, Chenwei Lin, Xian Xu

专题命中 文生图 :text-to-image(title,abstract);分类 cs.MM

AI总结 本文定义了T2I模型在应急场景下的EmergencyBias,构建评估框架发现其存在人口统计学与行为偏差,提出ActionAlign方法可减少行为差异并保留图像质量。

Comments 15 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27779 2026-07-31 cs.CV 新提交 79%

CXR-Retrieve: Compositional Text-to-Image Retrieval in Chest Radiography

CXR-Retrieve:胸部X线摄影中的组合式文本到图像检索

Tomer Erez, Moshe Kimhi, Chaim Baskin, Ehud Rivlin

机构 * Technion – Israel Institute of Technology(以色列理工学院) Ben-Gurion University of the Negev(内盖夫本-古里安大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 针对胸部X线检索的目标不匹配问题,提出CXR-Retrieve基准与标签感知对比微调方法,在双病理组合和否定查询的Precision@5上较CXR-CLIP分别提升8.5和22.0个百分点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24215 2026-07-28 cs.CV 新提交 79%

TreeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation

TreeAdapter:用于细粒度物种图像生成的分层分类法引导适配器组合

Yuze Sun, Zhongjie Duan, Yingda Chen

专题命中 文生图 :image generation(title);text-to-image(abstract);分类 cs.CV

AI总结 针对通用文本到图像模型在生成稀有生物物种图像时性能下降的问题,提出TreeAdapter框架,利用分层分类数据,通过在分类树节点附加轻量级适配器及两阶段训练范式,实现准确的细粒度图像生成,性能优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20155 2026-07-20 cs.CV cs.CL 版本更新 79%

NAMESAKES: Probing Identity Memorization in Text-to-Image Models

NAMESAKES: 探究文本到图像模型中的身份记忆

Morris Alper, Vasudha Varadarajan, Moran Yanuka, Angelina Wang, Hadar Averbuch-Elor

机构 * Carnegie Mellon University(卡内基梅隆大学) Tel Aviv University(特拉维夫大学) Cornell University(康奈尔大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 提出一种黑盒行为探针,无需参考照片或训练数据,即可区分文本到图像模型生成的图像是记忆还是虚构,并在NAMESAKES数据集上验证其有效性。

Comments Project page: https://namesakes-web.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08201 2026-07-10 cs.CV cs.AI 新提交 79%

TMI: Text-to-Image Meets Image-to-Image for Complementary Data Synthesis to Boost Long-Tailed Instance Segmentation

TMI:文本到图像与图像到图像结合用于互补数据合成以促进长尾实例分割

Hyeonseop Song, Seokhun Choi, Hoseok Do

机构 * LG Electronics(LG电子)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 研究针对大词汇量实例分割受长尾分布和类间模糊性限制的问题,提出结合文本到图像生成与上下文感知图像到图像编辑的混合管道,引入VRAIN编辑器,在LVIS基准测试中超越基线,有效提升分割性能。

Comments Accepted to ECCV 2026. The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24548 2026-07-02 cs.CV 新提交 79%

Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning

文本到图像模型是归纳主义的火鸡吗?一个用于因果推理的反事实基准

Jiayi Lei, Yuandong Pu, Xingyu Han, Rongpeng Zhu, Jing Xu, Jinyao Wang, Zijian Zhou, Bin Fu, Yuewen Cao, Yihao Liu, Hongsheng Li

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 提出反事实基准CF-World,通过三个递进层级测试T2I模型在违反现实先验规则下的图像生成能力,发现所有模型在反事实设置下性能急剧下降,原因是模型将世界知识与视觉外观编码为紧密耦合的模式。

Comments 10 pages, 7 figures. Project page: https://github.com/jylei16/CF-World.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30458 2026-06-30 cs.CV 79%

Cross-Resolution Semantic Transfer for Robust Text-to-Image Retrieval in Low-Resolution Surveillance

跨分辨率语义迁移用于低分辨率监控下的鲁棒文本-图像检索

Wenjie Qian, Bin Yang, Xiao Wang, Wenke Huang, Ling Mei, Xin Xu, Mang Ye

机构 * School of Computer Science and Technology, Wuhan University of Science and Technology(武汉科技大学计算机科学与技术学院) School of Computer Science, National Engineering Research Center for Multimedia Software, Wuhan University(武汉大学计算机学院,国家多媒体软件工程技术研究中心) Hubei Province Key Laboratory of Intelligent Information Processing and Real-time Industrial System, Wuhan University of Science and Technology(湖北省智能信息处理与实时工业系统重点实验室,武汉科技大学)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 针对低分辨率监控场景中文本-图像检索的可靠性崩溃和排序漂移问题,提出CLIP框架CRST,通过分辨率条件推理、文本引导精炼和跨分辨率邻域迁移,在三个数据集上平均提升超低分辨率Rank-1和mAP分别5.7%和5.3%。

Comments 10 pages,8 figures,conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27122 2026-06-30 cs.CV 79%

InterPartAbility: Phrase-Region Grounding for Interpretable Text-to-Image Person Re-Identification

InterPartAbility: 基于文本引导的部分匹配用于可解释的人员重识别

Shakeeb Murtaza, Aryan Shukla, Rajarshi Bhattacharya, Maguelonne Heritier, Eric Granger

机构 * LIVIA, Dept. of Systems Engineering, ETS Montreal, Canada(LIVIA系统工程系,蒙特利尔ÉTS学院,加拿大) Genetec Inc.(Genetec公司)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 本文提出InterPartAbility,通过显式部分匹配和短语-区域绑定提升TI-ReID的可解释性,引入PPIM模块实现概念级指导,生成 grounded 解释图谱,实验表明在CUHK-PEDES等基准上达到SOTA可解释性性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25297 2026-06-25 cs.CV 新提交 79%

Minimalist Preprocessing Approach for Image Synthesis Detection

图像合成检测的极简预处理方法

Hoai-Danh Vo, Trung-Nghia Le

机构 * University of Science, VNU-HCM(胡志明市国立大学理科大学) Vietnam National University, Ho Chi Minh City(胡志明市国立大学)

专题命中 文生图 :image synthesis(title);image generation(abstract);分类 cs.CV

AI总结 提出一种基于梯度计算相邻像素波动的轻量级预处理方法,作为高通滤波器突出关键特征,在低端设备上实现与先进技术相当的检测精度。

Comments SOICT 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06021 2026-06-24 cs.CV cs.AI cs.LG 79%

Improved Sub-Visible Particle Classification in Flow Imaging Microscopy via Generative AI-Based Image Synthesis

通过生成式AI图像合成改进流成像显微镜下的亚可见粒子分类

Utku Ozbulak, Michaela Cohrs, Hristo L. Svilenov, Joris Vankerschaver, Wesley De Neve

机构 * Center for Biosystems and Biotech Data Science(生物系统与生物技术数据科学中心) Ghent University Global Campus(根特大学全球校区) Department of Electronics and Information Systems(电子与信息系统系) Faculty of Pharmaceutical Sciences(药学系) Biopharmaceutical Technology, TUM School of Life Sciences(生物制药技术,技术大学生命科学学院) Department of Mathematics, Computer Science and Statistics(数学、计算机科学与统计学系)

专题命中 文生图 :image synthesis(title);diffusion(abstract);分类 cs.CV

AI总结 本文提出基于生成式AI的图像合成方法,解决流成像显微镜下亚可见粒子分类中的数据不平衡问题,通过生成高保真图像提升多类分类性能,公开模型和工具促进研究复现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23669 2026-06-23 cs.CV 新提交 79%

GeoFidelity-Bench: Evaluating Segment-Level Geographic Fidelity in Text-to-Image Street-View Generation

GeoFidelity-Bench:评估文本到图像街景生成中的片段级地理保真度

Kaizhen Tan, Hanzhe Hong, Siru Tao

机构 * Heinz College of Information Systems and Public Policy(信息系统与公共政策学院)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 提出GeoFidelity-Bench基准,通过参考面板排名测试文本到图像模型能否生成特定道路片段而非通用城市街景,发现添加街道和社区名称可提升检索准确率,但目标与最近邻片段间相似度差距近乎为零。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15967 2026-06-23 cs.CR cs.CV 版本更新 79%

When Safe Concepts Become Unsafe: Multi-Concept Compositional Vulnerabilities in Text-to-Image Models

TwoHamsters:文本到图像模型中多概念组合不安全性的基准测试

Chaoshuo Zhang, Yibo Liang, Mengke Tian, Chenhao Lin, Zhengyu Zhao, Le Yang, Chong Zhang, Yang Zhang, Qian Wang, Chao Shen

机构 * School of Cyber Science and Engineering, Xi'an Jiaotong University(西安交通大学计算机科学与工程学院) CISPA Helmholtz Center for Information Security(信息安全研究中心) School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院)

专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV

AI总结 本文提出TwoHamsters基准,通过17500个提示测试文本到图像模型在多概念组合不安全性的表现,揭示现有模型和防御机制在处理危险组合生成时的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏