arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OPAL:用于可解释图像分类的正交原型对齐学习

OPAL: Orthonormal Prototype Alignment Learning for Interpretable Image Classification

Ilán Carretero, Gustavo Jesús Angulo, Rocío del Amor, Valery Naranjo

arXiv 2608.30003首次发表:更新:

发表机构

Universitat Politècnica de València; Mines Paris, PSL University; Artikode Intelligence S.L.(瓦伦西亚理工大学; 巴黎矿业学院,巴黎文理研究大学; Artikode Intelligence S.L.)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出单阶段端到端框架OPAL,通过正交基锚定潜在空间、空间竞争机制实现可解释图像分类,性能优于同类模型,可提供细粒度视觉解释。

AI 中文摘要

基于原型部件的模型通过将输入区域与学习到的原型进行比较来提供可解释的预测,但现有方法存在复杂的多阶段训练流程,且严重依赖辅助正则化来防止原型坍塌。为克服这些局限,我们提出正交原型对齐学习(OPAL),这是一种简化可解释分类的单阶段端到端框架。该方法使用预定义的正交基锚定潜在空间,将每个类嵌入由固定部件原型张成的专用子空间中;为实现精确的部件定位,OPAL在特征图上施加空间竞争机制,该机制可分离稀疏的判别区域,引导每个原型在不同图像中始终关注同一语义概念。通过将分类构建为直接表示对齐任务,该方法消除了对辅助损失的需求。在细粒度基准上的大量实验表明,OPAL的性能优于不可解释的对应模型及当前最优的部件原型方法,还能通过明确揭示驱动每次预测的特定图像区域提供细粒度视觉解释。代码可在指定URL获取。

英文摘要

Prototypical part-based models provide explainable predictions by comparing input regions to learned prototypes. However, current approaches are burdened by complex, multi-stage training pipelines and heavily rely on auxiliary regularization to prevent prototype collapse. To overcome these limitations, we introduce Orthonormal Prototype Alignment Learning (OPAL), a single-stage, end-to-end framework that simplifies interpretable classification. Our approach anchors the latent space using predefined orthonormal bases, embedding each class within a dedicated subspace spanned by fixed part-prototypes. To achieve precise part localization, OPAL enforces spatial competition across feature maps. This mechanism isolates sparse, discriminative regions, directing each prototype to consistently attend to the same semantic concept across different images. By framing classification as a direct representation alignment task, our method eliminates the need for auxiliary losses. Extensive experiments on fine-grained benchmarks demonstrate that OPAL outperforms both its non-interpretable counterparts and state-of-the-art part-prototype methods, delivering granular visual explanations by explicitly revealing the specific image regions driving every prediction. Code is available at https://github.com/ilancarretero/OPAL.

CommentsAccepted at ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑