arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关系知识蒸馏使深度神经网络表征足够接近人类,从而无需监督即可对齐

Relational Knowledge Distillation Brings DNN Representations Close Enough to Humans to Be Aligned Without Supervision

Yuria Shimizu, Soh Takahashi, Takato Horii, Masafumi Oizumi

arXiv 2608.27877首次发表:更新:

发表机构

Graduate School of Arts and Sciences, The University of Tokyo; Graduate School of Engineering Science, The University of Osaka(东京大学文学部与科学研究科; 大阪大学工程科学研究科)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究采用Gromov-Wasserstein最优传输方法,发现用关系知识蒸馏微调预训练DNN,可在无监督下实现人类与DNN在单个物体层面的细粒度对齐,其改进源于更接近人类的全局结构。

AI 中文摘要

将深度神经网络(DNN)的内部表征与人类心理表征关联起来,对于将DNN用作人类视觉的计算模型至关重要。现有DNN表征与人类心理表征的相似性仍不足,而人类心理表征无法直接观测,因此通常通过对物体图像的大规模相似性判断来测量。缩小这一差距的自然方法是将人类表征的关系结构直接迁移到DNN中,已有研究报告称这可提高人类与DNN的表征相似性。然而,这种改进在更严格的评估下是否成立,尚未在两个方面得到验证:单个物体层面的细粒度对齐,以及对源自与训练数据无关的数据集的人类嵌入的泛化能力。在此,我们采用一种无监督比较方法——Gromov-Wasserstein最优传输(GWOT),该方法仅从内部距离结构估计人类与DNN的对应关系,从而测试细粒度对齐。我们还在与训练数据无重叠的精选概念测试集上评估泛化能力。结果显示,使用成熟的关系迁移方法Relational Knowledge Distillation(RKD)对预训练DNN进行微调,可使DNN在该测试集上的单个物体层面对齐中足够接近人类。我们还发现,这种改进由更接近人类的全局结构驱动,这反映在粗类别间的距离排序中,而局部人类与DNN的最近邻重叠率基本保持不变。这些发现表明,从人类进行的关系迁移使预训练DNN的全局结构足够接近人类结构,从而无需监督即可实现细粒度的人类与DNN对齐。

英文摘要

Linking the internal representations of deep neural networks (DNNs) to human mental representations is important for using DNNs as computational models of human vision. Existing DNN representations remain insufficiently similar to human mental representations, which are not directly observable and are therefore commonly measured through large-scale similarity judgments of object images. A natural approach to narrowing this gap is to directly transfer the relational structure of human representations into DNNs, and previous studies have reported improved human-DNN representational similarity. However, whether this improvement holds under stricter evaluation remains untested in two respects: fine-grained alignment at the individual-object level, and generalization to a human embedding derived from a dataset independent of the training data. Here, we employ an unsupervised comparison method, Gromov-Wasserstein optimal transport (GWOT), which estimates human-DNN correspondences from the internal distance structure alone and thereby tests fine-grained alignment. We further assess generalization on a curated test set of concepts non-overlapping with the training data. We show that fine-tuning pre-trained DNNs with Relational Knowledge Distillation (RKD), an established relational transfer method, brings DNNs close enough to humans to be aligned at the individual-object level on this test set. We also show that this improvement is driven by a more human-like global structure, as reflected in the ordering of distances among coarse categories, while the local human-DNN nearest-neighbor overlap rate remains largely unchanged. These findings indicate that relational transfer from humans brings the global structure of pre-trained DNNs close enough to the human structure to enable fine-grained human-DNN alignment without supervision.

Comments53 pages, 6 figures, 5 tables, including Supplementary Information (15 pages, 1 figure, 4 tables)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑