arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当组合失效时:人类识别AI生成图像中的缺陷

When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images

Ruoqi Hu, Chulin Zhao, Jiashuo Chang, Ramon Ruiz-Dolz, Hanhe Lin

arXiv 2608.25933首次发表:更新:

发表机构

Dundee International Institute of Central South University; Central South University; Faculty of Science, Engineering and Business, University of Dundee; University of Dundee(中南大学邓迪国际学院; 中南大学; 邓迪大学科学、工程与商学院; 邓迪大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文探究人类识别T2I模型生成图像组合缺陷的方法,构建CO-AID数据集,训练深度模型可预测并优化AI图像生成,验证了其可用性与有效性。

AI 中文摘要

Chulin Zhao与Ruoqi Hu对本文贡献等同。当前最先进的文本到图像(T2I)模型在提示涉及复杂组合因素(如多个实体、多个属性)时,会出现明显且系统性的缺陷。本文研究人类如何识别此类缺陷:首先从人物、手部、物体、场景四类中手动选取651张具有复杂组合特征的参考图像,通过手动编辑ChatGPT生成的提示,得到强调组合因素的提示;再将这些提示输入三个选定的T2I模型以生成AI图像,开展全面主观研究以识别缺陷。每张图像由29名参与者进行多标签评估,明确缺陷类型与位置。该研究产出了组合式AI生成图像缺陷(CO-AID)数据集,包含参考图像、提示、AI生成图像及缺陷位置与类型信息。实验结果显示,在CO-AID上训练深度模型,既可预测AI生成图像的缺陷,又能优化AI图像生成,证明了其可用性与有效性。数据库及补充材料可访问:this https URL。

英文摘要

*Chulin Zhao and Ruoqi Hu contributed equally to this work. State-of-the-art text-to-image (T2I) models exhibit pronounced and systematic defects when prompts involve intricate compositional factors such as multiple entities and multiple attributes. In this paper, we investigate how humans identify such defects. Specifically, we manually select 651 reference images from the four categories of people, hand, object, and scene that exhibit complex compositional characteristics, from which prompts emphasizing compositional factors are derived by manually editing ChatGPT-generated prompts. We then feed the prompts into three selected T2I models to generate AI images and conduct a comprehensive subjective study to identify their defects. For each image, 29 participants provide multi-label assessments specifying defect types and locations. The study yields the compositional AI-generated image defect (CO-AID) dataset, including reference images, prompts, AI-generated images, and information on defect locations and types. Experimental results show that training a deep model on CO-AID can both predict defects in AI-generated images and optimize AI image generation, demonstrating its usability and effectiveness. The database and supplementary materials are available at: https://github.com/Future-IQA/CO-AID .

Comments6 pages, accepted at IEEE MMSP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑