发表机构
L3i, University of La Rochelle; Itesoft(拉罗谢尔大学L3i研究所; Itesoft公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对文档图像分类中的LIME解释,发现标准超像素分割与文档结构不匹配,提出文档感知分割(基于OCR边界框和网格)能产生更稳定、更忠实的解释,并暴露识别码捷径偏差,强调分割应作为解释方法的一部分。
AI 中文摘要
事后解释方法被广泛用于检查图像分类器,但其可靠性依赖于常被视为实现细节的设计选择。我们针对文档图像分类中的LIME方法研究这一问题,重点关注定义被扰动可解释单元的分割步骤。基于图像的标准LIME通常依赖自然图像超像素,这些超像素与文档结构(如文本区域、版面块和识别码)对齐不佳。利用RVL-CDIP数据集,我们比较了Quickshift和SLIC与基于OCR边界框和规则网格的文档感知分割方法。结果表明,分割强烈影响解释的一致性、正确性和局部保真度。文档感知分割产生更稳定和忠实的解释,需要更少的扰动即可收敛,并暴露了基于文档识别码的捷径行为——这是RVL-CDIP中已知的偏差,而基于超像素的LIME常常掩盖该偏差。这些发现表明,可靠的事后解释需要领域感知的可解释表示,且分割应被视为解释方法的一部分,而非中性预处理。
英文摘要
Post-hoc explanation methods are widely used to inspect image classifiers, but their reliability depends on design choices that are often treated as implementation details. We study this issue for LIME on document image classification, focusing on the segmentation step that defines the interpretable units being perturbed. Standard image-based LIME typically relies on natural-image superpixels, which are poorly aligned with document structure such as text regions, layout blocks, and identification codes. Using RVL-CDIP, we compare Quickshift and SLIC with document-aware segmentations based on OCR bounding boxes and regular grids. Our results show that segmentation strongly affects explanation consistency, correctness, and local fidelity. Document-aware segmentations produce more stable and faithful explanations, require fewer perturbations to converge, and expose shortcut behaviour based on document identification codes, a known RVL-CDIP bias that superpixel-based LIME often obscures. These findings show that reliable post-hoc explanation requires domain-aware interpretable representations, and that segmentation should be treated as part of the explanation method rather than as neutral preprocessing.
Comments9 pages, 6 figures, 1 table, XAI Workshop @ IJCAI 2026