arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过重新审视基于扩散模型的跨域小样本目标检测数据生成实现无额外开销的数据增强

Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection

Zijian Zhuang, Yixiong Zou, Yuhua Li, Ruixuan Li

arXiv 2608.04394首次发表:更新:

发表机构

School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对跨域小样本目标检测的域差距与数据稀缺挑战,提出带定制噪声的选择性修复(SITN)方法,结合定制噪声生成与选择性修复模块合成有效数据,在6个CDFSOD和4个CDFSS数据集上达到新的最先进性能

AI 中文摘要

跨域小样本目标检测(CDFSOD)旨在利用稀缺的训练数据,将知识从数据丰富的上游通用域迁移到下游专业域,而显著的域差距和数据稀缺性使其成为尚未解决的挑战。为解决该问题,我们重新审视CDFSOD中一种自然但未被充分探索的方法:数据增强,即通过扩散模型直接合成数据以补充有限的训练样本。然而,由于存在较大的域差距,我们发现当前的扩散方法无法产生良好的结果,导致性能甚至低于使用原始图像的情况。为解决这些局限性,我们将域差距分为视觉差距和语义差距进行单独分析。对于视觉差距,我们发现扩散模型无法区分专业域中的噪声和有用信息,这可以通过添加弱化的噪声来缓解。对于语义差距,我们发现背景语义在域之间的差距远小于前景语义,我们可以通过背景修复来弥合该差距。基于上述分析,我们提出一种方法(带定制噪声的选择性修复,SITN),根据下游数据与通用域的不同差距,动态采用不同策略进行数据合成,包括用于添加定制噪声的生成模块和动态选择修复区域的选择模块。在6个CDFSOD数据集和4个跨域小样本分割(CDFSS)数据集上进行的大量实验验证,我们可以合成有用的数据,实现了新的最先进性能。我们的代码可在该https URL获取

英文摘要

Cross-Domain Few-Shot Object Detection (CDFSOD) aims to transfer knowledge from data-rich upstream generic domains to downstream expert domains using scarce training data, where the significant domain gap and data scarcity make it an unsolved challenge. To address this problem, we revisit a natural yet underexplored approach in CDFSOD: data augmentation, by directly synthesizing data through diffusion models to supplement limited training samples. However, due to large domain gaps, we find that current diffusion methods cannot produce good results, leading to performance even lower than using the original images. To address these limitations, we divide the domain gaps into visual gaps and semantic gaps for separate analysis. For the visual gap, we find that the diffusion model cannot distinguish noise from useful information on expert domains, which can be mitigated by adding weakened noise. For the semantic gap, we find that the background semantics shows much smaller gaps between domains than foreground semantics, and we can bridge this gap by background inpainting. Based on the above analysis, we propose a method (Selective Inpainting with Tailored Noise, SITN) to dynamically take different strategies for downstream data synthesis based on their different gaps from the general domain, including a Generation Module for adding tailored noise and a Selection Module to dynamically select the inpainting regions. Extensive experiments on 6 datasets of CDFSOD and 4 datasets of cross-domain few-shot segmentation (CDFSS) validate that we can synthesize helpful data, achieving new state-of-the-art performance. Our codes is available at https://github.com/zzzzj311-droid/Free-Lunch-SITN

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑