arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

集中而非不确定性:为何针对性合成数据无助于伪装目标检测

Concentration, Not Uncertainty: Why Targeted Synthetic Data Doesn't Help Camouflaged Object Detection

Akshat Dobhal, Sanjay Singh

arXiv 2610.09807首次发表:更新:

发表机构

Manipal Institute of Technology; Manipal Academy of Higher Education(马尼帕尔理工学院; 马尼帕尔高等教育学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究伪装目标检测中固定合成数据预算的分配策略,发现基于不确定性的针对性生成不优于随机分配,预算集中而非不确定性位置决定训练集效应,并揭示CHAMELEON数据集存在污染。

AI 中文摘要

伪装目标检测需要像素级精确的掩码,但获取此类标注既缓慢又昂贵,这使得合成训练图像成为一种有吸引力的替代方案。然而,在固定的生成预算下,尚不清楚应针对真实图像的哪些区域进行合成数据生成。我们研究了一种不确定性引导的生成策略,该策略对未标注的真实图像进行聚类,识别模型最不确定的聚类,将合成生成分配至这些聚类,并迭代地重新训练模型。在103次训练运行中,基于不确定性的针对性分配并未优于随机分配。五项独立对照进一步表明,这一零结果并非伪影:针对性训练集与随机训练集在可测量上存在差异,但该差异可通过集中生成预算而非不确定性集中的位置来解释,因为我们测试的每种集中规则都重现了该效应,并且在边界精度上,针对模型最确定的聚类也能产生同样效果。此外,我们发现在CHAMELEON数据集中存在大量数据污染,其76张图像中有50张与训练数据重复,尽管标准重叠检查报告重叠为零。综合来看,这些结果表明,在固定的合成数据预算下,预算集中而非基于不确定性的针对性分配,解释了所观察到的训练集效应。

英文摘要

Camouflaged object detection requires pixel-accurate masks, but obtaining such annotations is slow and costly, making synthetic training images an attractive alternative. Under a fixed generation budget, however, it remains unclear which real-image regions to target for synthetic data generation. We study an uncertainty-guided generation strategy that clusters the unlabelled real images, identifies clusters on which the model is least certain, allocates synthetic generation toward those clusters, and iteratively retrains the model. Across 103 training runs, uncertainty-based targeting does not outperform random allocation. Five independent controls further show that this null result is not an artifact: targeted training sets are measurably different from random sets, but the difference is explained by concentrating the generation budget rather than by where uncertainty is concentrated, as every concentration rule we test reproduces the effect and, on boundary accuracy, so does aiming at the clusters the model was most certain about. Separately, we find substantial data contamination in CHAMELEON, with 50 of its 76 images duplicated from training data despite the standard overlap check reporting zero overlap. Together, these results show that, under a fixed synthetic-data budget, budget concentration, not uncertainty-based targeting, accounts for the observed training-set effects.

Comments29 pages, 3 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑