arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29004cs.LGcs.CVcs.IR

面向检索与图卷积网络分类的上下文感知可解释表示

Context-Aware Interpretable Representations for Retrieval and Graph Convolutional Network Classification

发表机构圣保罗州立大学(UNESP)
查看机构详情
  • State University of São Paulo (UNESP)(圣保罗州立大学(UNESP))

机构由 AI 辅助整理,请以论文原文为准。

Thiago César Castilho Almeida, Gustavo Rosseto Letício, Vinicius Atsushi Sato Kawai, Daniel Carlos Guimarães Pedronette

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对视觉表示的几何与可解释性鸿沟,提出融合流形学习与可解释图嵌入的无监督框架,所获上下文感知表示兼具可解释性、降维能力,且在图像检索与GCN半监督分类任务中表现优异。

中文摘要 AI 辅助

过去几十年,视觉信息建模与表示领域取得了显著进展,这主要得益于卷积神经网络、基于Transformer的模型以及基础模型的支撑。尽管取得了这些进展,但相似性评估的本质和模型可解释性方面的关键挑战却被忽视了。一个核心问题是几何鸿沟:传统的成对度量无法捕捉数据集流形的内在几何结构。此外,可解释性鸿沟依然存在,因为表示通常与人类认知缺乏一致性。因此,如何在保持表示低维度和下游任务高有效性的同时为其提供可解释性,仍是一个未解决的挑战。本文提出了一种新颖的无监督框架,将流形学习策略与基于排序的可解释图嵌入相结合。该方法通过流形分析首先表征数据集的上下文信息,随后生成稀疏的、可自解释的嵌入,有效弥合了上述鸿沟。所提方法采用灵活的公式,支持不同的流形学习和表示学习策略。在各类数据集和特征上开展的大量实验评估表明,我们的上下文感知表示不仅提供内在可解释性并实现降维,还能在下游任务中维持或提升有效性,尤其适用于图像检索和使用图卷积网络(GCN)的半监督分类任务。

英文摘要

The advances in visual information modeling and representation during the last decades are remarkable, mainly supported by Convolutional Neural Networks, Transformer-based, and Foundation Models. Despite this progress, critical challenges regarding the nature of similarity assessment and model transparency have been neglected. A primary concern is the Geometric Gap, where traditional pairwise measures fail to capture the intrinsic geometry of the dataset manifold. Furthermore, the Interpretability Gap persists, as representations often lack alignment with human cognition. Therefore, how to provide interpretability to representations while maintaining low dimensionality and high effectiveness in downstream tasks remains an open challenge. In this paper, we propose a novel unsupervised framework that integrates Manifold Learning strategies with Rank-based Interpretable Graph Embeddings. Our approach effectively bridges these gaps by first characterizing the contextual information of the dataset through manifold analysis and subsequently generating sparse, self-explainable embeddings. The proposed approach employs a flexible formulation, allowing different Manifold Learning and Representation Learning strategies. Extensive experimental evaluation across diverse datasets and features demonstrates that our Context-Aware representations not only provide intrinsic interpretability and dimensionality reduction but also maintain or enhance effectiveness in downstream tasks, specifically in image retrieval and semi-supervised classification using Graph Convolutional Networks (GCNs).

补充信息

↑