arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86761 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3489 篇

2303.04001 2023-03-10 cs.CV cs.CL cs.GR cs.LG 73%

ELODIN: Naming Concepts in Embedding Spaces

Rodrigo Mello, Filipe Calegario, Geber Ramalho

专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV、cs.GR

Comments Added quantitative data, fixed formatting issues

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.03430 2023-02-21 cs.LG cs.AI cs.CL cs.CV cs.MM 73%

Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions

Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14794 2022-12-12 cs.CV cs.AI cs.LG cs.MM stat.ML 73%

Traditional Classification Neural Networks are Good Generators: They are Competitive with DDPMs and GANs

Guangrun Wang, Philip H. S. Torr

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.MM

Comments This paper has 29 pages with 22 figures, including rich supplementary information. Project page is at \url{https://classifier-as-generator.github.io/}

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08472 2022-04-20 cs.CV cs.GR cs.LG 73%

Simultaneous Multiple-Prompt Guided Generation Using Differentiable Optimal Transport

Yingtao Tian, Marco Cuturi, David Ha

专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV、cs.GR

Comments Accepted at ICCC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.09753 2021-10-20 cs.CV cs.CL cs.MM 73%

Unifying Multimodal Transformer for Bi-directional Image and Text Generation

Yupan Huang, Hongwei Xue, Bei Liu, Yutong Lu

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV、cs.MM

Comments ACM MM 2021 (Industrial Track). Code: https://github.com/researchmm/generate-it

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.09486 2021-06-18 cs.CV cs.GR 73%

Deep HDR Hallucination for Inverse Tone Mapping

Demetris Marnerides, Thomas Bashford-Rogers, Kurt Debattista

专题命中 文生图 :inpainting(abstract);image synthesis(abstract);分类 cs.CV、cs.GR

Journal ref Sensors 2021, 21, 4032

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.10299 2021-04-22 cs.GR cs.CV cs.LG cs.SD eess.AS 73%

Voice2Mesh: Cross-Modal 3D Face Model Generation from Voices

Cho-Ying Wu, Ke Xu, Chin-Cheng Hsu, Ulrich Neumann

专题命中 文生图 :image generation(abstract);image synthesis(abstract);分类 cs.CV、cs.GR

Comments Project page: https://choyingw.github.io/works/Voice2Mesh/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.00824 2019-04-02 cs.CV cs.GR stat.ML 73%

Training Object Detectors on Synthetic Images Containing Reflecting Materials

Sebastian Hartwig, Timo Ropinski

专题命中 文生图 :image generation(abstract);image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04734 2025-08-08 q-bio.QM cs.AI eess.IV 71%

Cross-Domain Image Synthesis: Generating H&E from Multiplex Biomarker Imaging

Jillur Rahman Saurav, Mohammad Sadegh Nasr, Jacob M. Luber

机构 * Department of CSE University of Texas at Arlington Arlington, Texas, USA(计算机科学与工程系 乌德勒支理工大学 美国德克萨斯州阿灵顿)

专题命中 文生图 :image synthesis(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02937 2025-08-06 cs.CY 71%

Documenting Patterns of Exoticism of Marginalized Populations within Text-to-Image Generators

Sourojit Ghosh, Sanjana Gautam, Pranav Venkit, Avijit Ghosh

专题命中 文生图 :text-to-image(title)

Comments Upcoming Publication, AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14285 2025-05-20 cs.CL 71%

Vulnerability of Text-to-Image Models to Prompt Template Stealing: A Differential Evolution Approach

Yurong Wu, Fangwen Mu, Qiuhong Zhang, Jinjing Zhao, Xinrun Xu, Lingrui Mei, Yang Wu, Lin Shi, Junjie Wang, Zhiming Ding, Yiwei Wang

机构 * Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of California at Merced(加州大学默塞德分校) University of Chinese Academy of Sciences(中国科学院大学) The University of Sydney(悉尼大学)

专题命中 文生图 :text-to-image(title)

Comments 14 pages,8 figures,4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13834 2024-12-19 cs.IR cs.AI cs.LG 71%

Maybe you are looking for CroQS: Cross-modal Query Suggestion for Text-to-Image Retrieval

Giacomo Pacini, Fabio Carrara, Nicola Messina, Nicola Tonellotto, Giuseppe Amato, Fabrizio Falchi

专题命中 文生图 :text-to-image(title)

Comments 15 pages, 5 figures. To be published as full paper in the Proceedings of the European Conference on Information Retrieval (ECIR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16593 2024-07-23 cs.HC 71%

Investigating the Design Considerations for Integrating Text-to-Image Generative AI within Augmented Reality Environments

Yongquan Hu, Dawen Zhang, Mingyue Yuan, Kaiqi Xian, Don Samitha Elvitigala, June Kim, Gelareh Mohammadi, Zhenchang Xing, Xiwei Xu, Aaron Quigley

专题命中 文生图 :text-to-image(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.01593 2017-01-09 physics.geo-ph 71%

Image synthesis with graph cuts: a fast model proposal mechanism in probabilistic inversion

T. Zahner, T. Lochbühler, G. Mariethoz, N. Linde

专题命中 文生图 :image synthesis(title)

Journal ref Geophysical Journal International, 204, 1179-1190 (2016)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20382 2026-08-24 cs.CL cs.CV 新提交 70%

Decoupled Vision-Language System for Multimodal Understanding and Generation

用于多模态理解与生成的解耦式视觉-语言系统

Yifan Xu, Baochen Xiong, Xiaoshan Yang, Donglin Di, Yaowei Wang, Changsheng Xu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) Peng Cheng Laboratory(鹏城实验室) Li Auto(理想汽车)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 本研究提出多模态大语言模型Libra的解耦式视觉-语言架构,通过开关注意力与FFN模块实现自模态与跨模态解耦,在理解与生成任务上均取得优异性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13112 2026-08-14 cs.CV 新提交 70%

Towards Physics-Faithful Generation of Scientific Diagrams

面向物理保真的科学图表生成

Minghui Zhang, Jinxin Shi, Yifan Chang, Liangliang Zhao, Yuandong Pu, Qian Yu, Ming Hu, Hanxiao Zhang, Yun Gu, Yirong Chen, Yu Qiao, Bo Zhang, Xiangchao Yan, Bin Fu, Yihao Liu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 针对现有文本到图像生成模型生成的科学图表存在物理错误的问题,提出Princigram生成器,通过结构化物理思维链实现物理保真的科学图表生成,在相关基准上验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05210 2026-08-07 cs.CV cs.AI cs.CR 新提交 70%

Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation

无辜的面板,仇恨的故事:评估与检测多轮视觉故事生成中的仇恨意图

Ye Leng, Junjie Chu, Yiting Qu, Mingjie Li, Yun Shen, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心) Hewlett Packard Enterprise(惠普企业)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 本研究针对多轮视觉故事生成的组级仇恨意图问题,构建了专用评估数据集,发现现有审核系统存在漏检,提出互补防御方法,强调安全需适配视觉叙事的发展。

Comments 16 pages, 5 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24898 2026-07-29 cs.CV cs.AI 新提交 70%

Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed

危害并非普遍存在:迫切需要针对特定社区的毒性检测

Xinnuo Xu, Anja Thieme, Daniela Massiceti, Ioana Tanase, Rita Marques, Melanie Fernandez Pradier, Martin Grayson, Camilla Longden, Cecily Morrison

机构 * Microsoft Research Cambridge, UK(英国剑桥微软研究院) Microsoft Paris, France(法国巴黎微软)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 研究指出文本到图像生成的通用毒性检测方法不能保护边缘化社区,主张特定社区毒性检测(CTD)。通过与专家合作制定指南,利用图像数据集实验表明现有模型表现不佳,基于提示的方法和参数高效微调可提升性能,但CTD性能仍远低于通用检测,需持续研究。

Comments 18 pages, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23271 2026-07-28 cs.CV cs.AI 新提交 70%

What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features

CLIP 所知道但无法表达的:从冻结的中间特征中恢复否定信息

Chen-Yi Lu, Yueh-Shao Chen, Somali Chaterji

机构 * Purdue University(普渡大学)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 对比视觉语言模型如CLIP对否定不敏感,研究发现是表征坍缩所致。提出轻量级事后校正系统PeakPatch,在不改变预训练权重下恢复否定信号,通过特定网络提取信号、预测偏差向量等,经实验验证其有效性及泛化性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17951 2026-07-24 cs.CV 版本更新 70%

SuperFlow: Training Flow Matching Models with RL on the Fly

SuperFlow: 通过实时强化学习训练流匹配模型

Kaijie Chen, Zhiyang Xu, Ying Shen, Zihao Lin, Yuguang Yao, Lifu Huang

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 SuperFlow通过实时强化学习优化流模型训练,提升文本-图像生成效率和质量。

Comments This article is withdrawn because it was submitted to arXiv without obtaining the consent of all listed authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19558 2026-07-14 cs.CR cs.AI cs.CV cs.LG 版本更新 70%

SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation

SPQR:良性模型适应下安全对齐的多维基准测试

Mohammed Talha Alam, Nada Saadi, Fahad Shamshad, Nils Lukas, Karthik Nandakumar, Fahkri Karray, Samuele Poppi

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) University of Waterloo(滑铁卢大学) Michigan State University(密歇根州立大学)

专题命中 文生图 :text-to-image(abstract);diffusion(abstract);分类 cs.CV

AI总结 研究文本到图像扩散模型安全对齐在良性微调下的稳定性,引入SPQR基准测试,通过单评分指标提供统一框架,经多方面分析确定安全对齐失败情况,为T2I安全对齐技术提供简洁全面的基准。

Comments 34 pages, 9 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12886 2026-07-10 cs.CV cs.AI 新提交 70%

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement

交错思维中的模态隔离桥接:通过逐步强化监督模态转换

Tingyu Li, Le Zhou, Siyuan Li, Yujun Wu, Xinglong Xu, Jingxuan Wei, Conghui He, Cheng Tan

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiaotong University(上海交通大学) Zhejiang University(浙江大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 提出MoTiF框架,通过反射式SFT和Flow-GRPO优化模态转换保真度,解决交错思维中图像与文本脱节的模态隔离问题,提升跨模态一致性和任务准确性。

Comments 22 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06839 2026-07-09 cs.LG cs.CV 新提交 70%

LEMUR 2: Unlocking Neural Network Diversity for AI

LEMUR 2:释放人工智能的神经网络多样性

Tolgay Atinc Uzun, Waleed Khalid, Saif U Din, Sai Revanth Mulukuledu, Akashdeep Singh, Chandini Vysyaraju, Raghuvir Duvvuri, Avi Goyal, Yashkumar Rajeshbhai Lukhi, Muhammad A. Hussain, Krunal Jesani, Usha Shrestha, Yash Mittal, Roman Kochnev, Pritam Kadam, Mohsin Ikram, Harsh R. Moradiya, Alice Arslanian, Dmitry Ignatov, Radu Timofte

机构 * University of Würzburg(维尔茨堡大学)

专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV

AI总结 LEMUR 2旨在释放神经网络多样性,通过统一生成、评估和部署管道,利用多种方式生成超14000个不同架构及大量训练记录,采用特定管道进行自动部署和延迟基准测试,涵盖多模态任务,为LLM微调提供数据基础,推动相关新兴范式发展。

Comments 10 pages, 9 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07778 2026-06-30 cs.CV 70%

Distribution Matching Variational AutoEncoder

分布匹配变分自编码器

Sen Ye, Jianning Pei, Mengde Xu, Shuyang Gu, Chunyu Wang, Liwei Wang, Han Hu

专题命中 文生图 :diffusion(abstract);image synthesis(abstract);分类 cs.CV

AI总结 本文提出DMVAE,通过分布匹配约束将编码器的潜在分布与任意参考分布对齐,发现基于自监督学习的分布在图像生成中表现优异。

Comments ICML2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21081 2026-06-19 cs.CV 版本更新 70%

Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought

思维形状:通过视觉思维链进行渐进式物体组装

Yu Huo, Siyu Zhang, Kun Zeng, Haoyue Liu, Owen Lee, Junlin Chen, Yuquan Lu, Yifu Guo, Yaodong Liang, Xiaoying Tang

机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)科学与工程学院) School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院) Sun Yat-sen University(中山大学) The Hong Kong University of Science and Technology, Guangzhou(香港科学与技术大学(广州)) Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen)(深圳未来网络智能研究所(FNii-Shenzhen)) Guangdong Provincial Key Laboratory of Future Networks of Intelligence, CUHK(SZ)(广东省未来网络智能重点实验室,CUHK(SZ))

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 提出Shape-of-Thought (SoT)框架,通过视觉思维链在渲染2D域中逐步组装形状,解决文本到图像生成中的组合结构约束问题,在组件计数和结构拓扑上显著优于直接生成。

Comments ICML2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14798 2026-06-18 cs.LG cs.CV 版本更新 70%

RUB: Evaluating Residual Knowledge in Unlearned Models

RUB: 评估未学习模型中的残留知识

Hao Xuan, Xingyu Li

机构 * Electrical and Computer Engineering University of Alberta(电气与计算机工程大学阿尔伯塔大学)

专题命中 文生图 :text-to-image(abstract);image synthesis(abstract);分类 cs.CV

AI总结 提出鲁棒未学习原则及统一基准RUB,通过未学习映射攻击(UMA)检测残留信息,揭示现有方法在对抗评估下的脆弱性。

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2026, pages 8550-8559

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10496 2026-06-16 cs.CV 版本更新 70%

CheXGenBench: A Unified Benchmark For Fidelity, Privacy and Utility of Synthetic Chest Radiographs

CheXGenBench:合成胸片保真度、隐私和实用性的统一基准

Raman Dutt, Pedro Sanchez, Yongchen Yao, Steven McDonagh, Sotirios A. Tsaftaris, Timothy Hospedales

机构 * University of Edinburgh(爱丁堡大学) Samsung AI Center, Cambridge(剑桥三星AI中心)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 提出CheXGenBench,首个统一评估框架,同时衡量合成胸片生成模型的保真度、隐私风险和下游实用性,涵盖11种前沿T2I模型,揭示当前模型在长尾分布、隐私风险和下游多模态任务中的局限。

Comments Published in Transactions of Machine Learning Research (06/2026)

Journal ref Transactions on Machine Learning Research (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13676 2026-06-12 cs.CV 新提交 70%

Modality Forcing for Scalable Spatial Generation

模态强制实现可扩展的空间生成

Bardienus Pieter Duisterhof, Deva Ramanan, Jeffrey Ichnowski, Justin Johnson, Keunhong Park

机构 * Carnegie Mellon University(卡内基梅隆大学) World Labs

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 提出Modality Forcing方法,通过为每个模态分配独立噪声水平,实现单DiT的联合图像-深度生成,利用稀疏深度数据训练,继承T2I预训练的可扩展性,在深度估计上取得竞争性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13288 2026-06-12 cs.CV cs.AI cs.CL 新提交 70%

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality

跨模态掩码组合概念建模以增强视觉-语言组合性

Wei Li, Zhen Huang, Xinmei Tian

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学,教育部脑启发智能感知与认知重点实验室) Independent Researcher(独立研究员)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 提出MACCO框架,通过掩码一个模态的组合概念并从另一模态完整上下文重建,增强视觉-语言模型的组合理解能力,在五个基准上显著提升。

Comments Accepted to ACL 2026 Main Conference, 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11188 2026-06-10 cs.CV 新提交 70%

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations

ARM: 一种具有统一离散表示的自回归大型多模态模型

Junke Wang, Xiao Wang, Jiacheng Pan, Xuefeng Hu, Feng Li, Jingxiang Sun, Chaorui Deng, Zilong Chen, Yunpeng Chen, Kaibin Tian, Matthew Gwilliam, Hao Chen, Danhui Guan, Kun Xu, Weilin Huang, Zuxuan Wu, Haoqi Fan, Yu-Gang Jiang, Zhenheng Yang

机构 * Shanghai Key Lab of Intelligent Information Processing, Fudan University(复旦大学上海智能信息处理重点实验室) School of Computer Science, Fudan University(复旦大学计算机科学技术学院) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) Youtu Lab, Tencent(腾讯优图实验室) Meta AI Shanghai AI Laboratory(上海人工智能实验室)

专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV

AI总结 提出ARM模型,通过离散语义视觉分词器将图像映射为紧凑token序列,结合自回归建模和强化学习,统一实现图像理解、生成和编辑,并提升任务性能与跨任务协同。

Comments technical report

详情

展开后加载摘要…

URL PDF HTML 收藏