CRD-CGAN: Category-Consistent and Relativistic Constraints for Diverse Text-to-Image Generation
专题命中 文生图 :image generation(title);text-to-image(title);分类 cs.CV
视觉与机器人
图像生成、文生图、图像编辑、扩散模型和可控生成。
专题命中 文生图 :image generation(title);text-to-image(title);分类 cs.CV
专题命中 文生图 :text-to-image(title);image synthesis(title);分类 cs.CV
Comments CVPR2018 Spotlight
MASCOT:面向复合属性文本到图像检索的模型感知子模覆盖
专题命中 文生图 :text-to-image(title,abstract);分类 cs.CV、cs.MM
AI总结 针对现有流形重排序方法在复合属性文本到图像检索的多样性降低任务中早期排名召回率大幅下降的问题,提出MASCOT方法,在PixelProse数据集复合约束任务中表现优于MS-DPP,尤其在排名1后召回率上优势显著。
Comments 21 pages, 4 figures. Accepted at ACM Multimedia 2026 (MM '26), Rio de Janeiro, Brazil. Extended version with full appendices
在文本到图像模型中实现公平性:对偏见、公平性审计及缓解策略的综述
机构 * AImotion Bavaria, Technische Hochschule Ingolstadt(AImotion巴伐利亚、因戈尔施塔特技术大学) ; School of Transformation and Sustainability, Catholic University of Eichstätt-Ingolstadt(转型与可持续性学院,埃施塔特-因戈尔施塔特天主教大学) ; Chair of Economic and Social Ethics, University of Hohenheim(经济与社会伦理系,霍亨海姆大学)
专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract);分类 cs.CV、cs.MM
AI总结 本文综述了文本到图像模型中的公平性研究,分析了偏见类型和公平性概念的分类,指出现有研究在目标公平与阈值公平之间的差距,并提出新的公平性操作框架。
Comments ICLR 2026 Algorithmic Fairness Across Alignment Procedures and Agentic Systems (AFAA) Workshop, reviews can be found at: https://openreview.net/forum?id=8DOkyBGWwP
Unify-Agent:一种用于世界 grounded 图像合成的统一多模态代理
机构 * University of California, Los Angeles(加州大学洛杉矶分校) ; Tencent Hunyuan(腾讯混元) ; The Chinese University of Hong Kong(香港中文大学) ; The Hong Kong University of Science and Technology(香港科技大学)
专题命中 文生图 :image synthesis(title,abstract);image generation(abstract);分类 cs.CV、cs.MM
AI总结 本文提出Unify-Agent,一种统一多模态代理,用于世界 grounded 图像合成,通过代理流程提升图像生成质量,结合事实IP基准测试,展示其在多种任务中的优越表现。
Comments Project Page: https://github.com/shawn0728/Unify-Agent
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Tsinghua University(清华大学) ; The Chinese University of Hong Kong(香港中文大学) ; CPII under InnoHK(创新香港科技促进会) ; MThreads AI
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);分类 cs.CV、cs.MM
Comments code: https://github.com/SunzeY/X-Prompt
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);分类 cs.CV、cs.MM
Comments 10 pages, 10 figures, 3 tables
专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);分类 cs.CV、cs.GR
Comments This work was presented at The European Conference on Computer Vision (ECCV) 2024 Workshop "Fairness and ethics towards transparent AI: facing the chalLEnge through model Debiasing" (FAILED), Milano, Italy, on September 29, 2024, https://failed-workshop-eccv-2024.github.io
专题命中 文生图 :image synthesis(title,abstract);image generation(abstract);分类 cs.CV、cs.MM
Comments Accepted to ACL2024 main
专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract);分类 cs.CV、cs.GR
Comments Project page at https://lcm-lookahead.github.io/
专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract);分类 cs.CV、cs.GR
Comments In International Conference on Learning Representations 12 (ICLR 2024) [79 pages, 54 figures, 7 tables]
专题命中 文生图 :image synthesis(title,abstract);text-to-image(abstract);分类 cs.CV、cs.GR
专题命中 文生图 :image synthesis(title,abstract);text-to-image(abstract);分类 cs.CV、cs.MM
专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);分类 cs.CV、cs.GR
Comments Project page at https://datencoder.github.io
专题命中 文生图 :text-to-image(title,abstract);image generation(abstract);分类 cs.CV、cs.GR
专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract);分类 cs.CV、cs.GR
Comments Project page at https://tuning-encoder.github.io/
专题命中 文生图 :image synthesis(title,abstract);image editing(abstract);分类 cs.CV、cs.MM
Comments ECCV 2022
Journal ref ECCV 2022
专题命中 文生图 :image synthesis(title,abstract);image generation(abstract);分类 cs.CV、cs.GR
Comments 26 pages, 7 figures, 1 table
Journal ref Computational Visual Media 2021
专题命中 文生图 :image generation(title,abstract);image synthesis(abstract);分类 cs.CV、cs.GR
Comments NeurIPS 2018. Code: https://github.com/junyanz/VON Website: http://von.csail.mit.edu/
专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract);分类 cs.CV
Comments Accepted by ACM CCS 2024. Please cite this paper as "Xinfeng Li, Yuchen Yang, Jiangyi Deng, Chen Yan, Yanjiao Chen, Xiaoyu Ji, Wenyuan Xu. SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models. In Proceedings of ACM Conference on Computer and Communications Security (CCS), 2024."
专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract);分类 cs.CV
Comments Github repo can be found at: https://github.com/harveymannering/Text-to-Image-Bias
X-MULTI:基于VLM的成像因子解耦用于因子感知图像合成
机构 * University of Siegen(锡根大学) ; ETH Zürich(苏黎世联邦理工学院) ; Bosch Research(博世研究中心) ; INSAIT(INSAIT(保加利亚的智能与数据科学研究所)) ; Sofia University “St. Kliment Ohridski”(索非亚大学“圣·克利门特·奥赫里德斯基”)
专题命中 文生图 :image synthesis(title);image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 该研究针对文本到图像生成的成像因子解耦问题,提出基于VLM的X-MULTI方法及改进的I-FAA指标,提升了新颖因子组合的因子对齐效果与评估的稳健性。
Comments Accepted to the MUCG Workshop at ECCV 2026
科学图像合成:基准测试、方法论与下游应用
机构 * Shanghai Jiao Tong University(上海交通大学) ; OpenDataLab, Shanghai Artificial Intelligence Laboratory(OpenDataLab,上海人工智能实验室) ; The University of Hong Kong(香港大学) ; Peking University(北京大学)
专题命中 文生图 :image synthesis(title,abstract);text-to-image(abstract);分类 cs.CV
AI总结 本文提出ImgCoder框架和SciGenBench基准,通过逻辑驱动方法提升科学图像生成的结构精度,并展示微调LMMs在科学图像上的效果,验证了高保真合成在多模态推理中的潜力。
GenRouter:面向智能体图像生成的统一工作流路由
机构 * HKUST(GZ)(香港科技大学(广州)) ; HKUST(香港科技大学) ; SUSTech(南方科技大学) ; ZODA ; CUHK(香港中文大学)
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);分类 cs.CV
AI总结 GenRouter是首个智能体图像生成的统一工作流路由框架,通过需求分析、经验匹配、帕累托过滤实现自适应路由,可大幅降低计算开销与延迟并提升视觉对齐效果。
Comments Code: https://github.com/EnVision-Research/GenRouter
CoCA:基于强化学习的文本到图像(T2I)扩散模型微调中的免费步骤级奖励
机构 * Huazhong University of Science and Technology(华中科技大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 文生图 :diffusion(title,abstract);text-to-image(abstract);分类 cs.CV
AI总结 针对RL驱动T2I扩散模型微调的奖励稀疏问题,提出CoCA信用分配框架,通过余弦相似度变化分配密集奖励,提升样本效率与泛化性且不损害原最优策略。
面向表达性与忠实性的音频到图像生成:一个统一的多模态数据集与合成框架
机构 * University of Science and Technology of China(中国科学技术大学) ; China Telecom (TeleAI)(中国电信(电信人工智能研究院)) ; Northwest Polytechnical University(西北工业大学)
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);分类 cs.CV
AI总结 针对音频到图像生成受限于传统数据集的问题,提出A2I-Set数据集与AudioCanvas模型,实现了更优的跨模态对齐与视觉表达性。
Comments 23 pages, 16 figures
ToolArtist:用于智能体图像生成的工具使用统一多模态模型
机构 * RUC(中国人民大学) ; HKUST(GZ)(香港科技大学(广州)) ; NUS(新加坡国立大学) ; UCD(加州大学戴维斯分校)
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);分类 cs.CV
AI总结 本文提出ToolArtist,一种全智能体图像生成模型,通过后训练UMM实现,采用RAD-GRPO方法优化,将完整开放世界图像生成过程置于智能体控制下,性能优于部分控制方案。
探究叙事图像生成中的社会偏见
机构 * KAIST(韩国科学技术院)
专题命中 文生图 :image generation(title,abstract);text-to-image(abstract);分类 cs.CV
AI总结 本研究将偏见评估框架BBG适配图像生成,对比6种T2I模型在照片、分镜、漫画生成中的偏见表现,发现叙事视觉形式中偏见更易显现,强调需用多样视觉形式评估T2I系统。
Comments Accepted to GenAI4World Workshop at COLM 2026
当漂亮并不有用:调查现代文本到图像模型为何无法作为可靠的训练数据生成器
机构 * RPTU University Kaiserslautern-Landau(凯撒斯劳滕-兰道大学) ; German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract);分类 cs.CV
AI总结 研究发现现代文本到图像模型在生成合成数据时性能下降,揭示其生成的数据分布狭窄且缺乏多样性,挑战了生成现实感进步即意味着数据现实感进步的假设。
Comments Accepted to CVPR26, Project Page: https://bill2462.github.io/When-Pretty-Isn-t-Useful-page/, Code: https://github.com/Bill2462/When-Pretty-Isn-t-Useful-codebase
AEGIS:一种针对文本到图像模型中视觉同义词越狱的机制引导防御
机构 * Fudan University(复旦大学) ; Alibaba Group(阿里巴巴集团) ; Shanghai Pudong Research Institute of Cryptology(上海浦东密码研究所) ; Engineering Research Center of Cyber Security Auditing and Monitoring, Ministry of Education(教育部网络安全审计与监测工程研究中心)
专题命中 文生图 :text-to-image(title,abstract);diffusion(abstract);分类 cs.CV
AI总结 研究针对文本到图像模型中视觉同义词越狱问题,提出AEGIS防御机制,通过动态追踪不安全语义生成过程,利用稀疏语义注入注意力头,在推理时仅对易受攻击头部应用相似性感知排斥,提升了模型安全与效用。