arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13729cs.CV

合成数据生成在数据稀缺的专业领域中的局限性

Limitations of Synthetic Data Generation in Specialized Data-Scarce Domains

发表机构宾夕法尼亚大学 · 香港理工大学
查看机构详情
  • University of Pennsylvania(宾夕法尼亚大学)
  • The Hong Kong Polytechnic University (PolyU)(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Edward Zhang, Marcel Hussing, Tanay Tandon, Shenbagaraj Kannapiran, Jason Hughes, Youkang Wang, Joshua Caswell, Agelos Kratimenos, Yi Fan Li, Milan Manoj, Ethan… 展开作者

Edward Zhang, Marcel Hussing, Tanay Tandon, Shenbagaraj Kannapiran, Jason Hughes, Youkang Wang, Joshua Caswell, Agelos Kratimenos, Yi Fan Li, Milan Manoj, Ethan Sanchez, Sumukh Shrote, Camillo Jose Taylor, Daniel A. Hashimoto, Eric Eaton

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对数据稀缺的专业视觉领域,评估两类生成式稀疏数据扩展方法的分类性能,发现其无法持续优于非生成式基线,且存在记忆坍缩等失败模式。

中文摘要 AI 辅助

基于扩散的生成模型的进展推动了合成图像生成技术的应用,以缓解视觉任务中的数据稀缺问题。尽管该策略在ImageNet等自然图像基准中展现出潜力,但其在稀疏、高方差的真实领域中的有效性仍不明确。本研究聚焦于图像与常见图像数据集差异显著且额外数据获取成本高昂的领域,针对非生成式数据增强基线,评估两类生成式稀疏数据扩展方法(分布建模与样本扰动)所带来的下游分类器性能提升。在使用按受试者划分的训练-验证集的5项创伤分类任务中,无一生成式方法能持续优于强大的非生成式基线。特征空间分析揭示了反复出现的失败模式:记忆或坍缩、分布漂移,以及生成视觉上合理但简化的规范实例,这些实例比真实数据更易分类。

英文摘要

Advances in diffusion-based generative models have motivated the use of synthetic image generation to alleviate data scarcity in vision tasks. While this strategy has shown promise in natural image benchmarks such as ImageNet, its effectiveness in sparse, high-variance real-world domains remains unclear. In this work, we focus on domains where images differ substantially from common image datasets and additional data are expensive to obtain. Against non-generative data augmentation baselines, we evaluate the downstream classifier performance improvements yielded by two schools of generative sparse data extension: distribution modeling and sample perturbation. Across five trauma classification tasks using subject-wise train--validation splits, no generative approach consistently outperforms a strong non-generative baseline. Feature-space analysis reveals recurring failure modes: memorization or collapse, distributional drift, and generation of visually plausible but simplified canonical instances that are easier to classify than real data.

↑