arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.02637cs.LGcs.AI

通过同构-异构分割对合成图像进行生成后筛选

Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting

Disheng Liu, Tuo Liang, Chaoda Song, Yu Yin

首次发表
浏览论文内容

中文总结 AI 辅助

研究给定固定合成图像池,能否通过选择信息子集提高下游效用。基于生成器倾向,将真实类别分为同构和异构子集,用保真度-多样性标准评分,该方法与生成器无关且无需再训练,效果良好。

中文摘要 AI 辅助

近期生成模型能产生高质量合成图像,为数据饥渴模型提供可扩展训练数据。现有利用潜力的方法有局限性。本文研究给定固定生成图像池,能否仅通过选择信息子集提高下游效用,答案是肯定的。基于对生成器结构偏差的洞察,提出方法并验证其有效性,表明生成后筛选是提高合成数据效用的补充机制。

英文摘要

Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models. Existing approaches to exploiting this potential typically involve 1) training or fine-tuning generators, or 2) using lightweight post-hoc adaptation like prompt engineering or inference-time guidance, making them generator-specific and expertise-intensive. We study a complementary question: given a fixed pool of generated images, can downstream utility be improved purely by selecting an informative subset? The answer is yes. We show that effective selection must counter a structural bias of modern generators: they tend to over-produce canonical modes of each class while underrepresenting intra-class variation. Building on this insight, we split each real class into a canonical Homogeneous (HO) subset and a non-redundant Heterogeneous (HE) subset, then score synthetic images by a fidelity-diversity criterion that rewards semantic alignment while penalizing canonical redundancy. The method is generator-agnostic and requires no retraining. Across multiple benchmarks, it consistently outperforms state-of-the-art data selection baselines and matches the real-data performance with up to 40% fewer synthetic samples. The same criterion remains effective when applied on top of stronger task-tuned generators, with gains on both classification and segmentation tasks. Post-generation selection is therefore not a substitute for better generators, but a complementary mechanism for improving the utility of synthetic data.

发表机构

  • Department of Computer and Data Sciences(计算机与数据科学系)

机构由 AI 辅助整理,请以论文原文为准。

↑