发表机构
University of Dubai(迪拜大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文量化了基于Patch的高光谱图像分类中随机采样导致的空间重叠数据泄漏,提出OP和AOR度量,实验表明3D-CNN在随机采样下OA达96.17%,非随机下降至55.20%,证明随机评估会显著高估性能。
AI 中文摘要
基于Patch的学习通过利用局部光谱-空间信息提高了高光谱图像(HSI)分类的性能,但从同一图像中随机划分训练集和测试集可能导致空间Patch重叠,从而引发数据泄漏和乐观的性能估计。本文研究了基于Patch的HSI分类中同类训练-测试空间重叠问题,使用两种度量:重叠百分比(OP),用于量化重叠的测试Patch像素的全局数量;以及平均重叠比率(AOR),用于衡量受影响的测试Patch之间的局部严重程度。在Pavia University数据集上的实验比较了随机和非随机空间采样,使用了SVM、MLP、2D-CNN、3D-CNN、ViT和MorpMamba。结果表明,深度基于Patch的模型在随机采样下实现了高精度,其中3D-CNN达到了96.17%的总体精度(OA),但在非随机空间采样下大幅下降,3D-CNN降至55.20%,ViT和2D-CNN分别下降了40.71和38.81个百分点(PP)。Patch大小分析进一步表明,将Patch大小从5x5增加到19x19,随机采样的重叠百分比从23.28%上升到77.02%。这些发现表明,随机的基于Patch的评估可能显著夸大分类性能,尤其是对于强烈利用空间上下文的模型。本文相关的代码可在以下网址获取:this https URL。
英文摘要
Patch-based learning improves hyperspectral image (HSI) classification by exploiting local spectral-spatial information, but random train-test sampling from the same image can cause spatial patch overlap, leading to data leakage and optimistic performance estimates. This paper investigates same-class train-test spatial overlap in patch-based HSI classification using two measures: overlap percentage (OP), which quantifies the global amount of overlapped testing patch pixels, and average overlap ratio (AOR), which measures the local severity among affected testing patches. Experiments on the Pavia University dataset compare random and non-random spatial sampling using SVM, MLP, 2D-CNN, 3D-CNN, ViT, and MorpMamba. The results show that deep patch-based models achieve high accuracy under random sampling, with 3D-CNN reaching 96.17% Overall Accuracy (OA), but drop substantially under non-random spatial sampling, where 3D-CNN decreases to 55.20% and ViT and 2D-CNN drop by 40.71 and 38.81 percentage points (PP), respectively. Patch-size analysis further shows that increasing the patch size from 5x5 to 19x19 raises the random-sampling overlap percentage from 23.28% to 77.02%. These findings demonstrate that random patch-based evaluation can substantially inflate classification performance, especially for models that strongly exploit spatial context. The code associated with this paper is available at: https://github.com/mqalkhatib/Data_Leakage_in_HSI_Classification.
Commentspaper accepted for presentation at IEEE-WHISPERS