arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可解码但不可访问:对解耦皮肤病变表征的基于距离的可靠性估计的审计

Decodable but Not Accessible: Auditing Distance-Based Reliability Estimation on Disentangled Skin-Lesion Representations

Duc-Vinh Tran

arXiv 2608.11267首次发表:更新:

发表机构

School of Information and Communication Technology, Hanoi University of Science and Technology(河内科技大学信息与通信技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究审计了基于距离的可靠性估计假设,发现解耦皮肤病变表征的几何结构变化未提升可靠性估计,信息可解码但对非探针估计器不可用。

AI 中文摘要

基于距离的可靠性估计假设表征的几何结构反映其可信度,但这一假设在直接重塑几何结构的训练干预下很少被测试。我们使用解耦剂量响应阶梯,在领域对抗表征学习下审计该假设。三组检查点家族共享相同架构和16维表征,仅正交性强度不同(λ=0、1、5)。表征几何结构随解耦强度发生显著变化:条件数偏移了两个数量级(Kendall tau=0.84,精确p=2.8e-5)。这一变化并未伴随可靠性估计的提升:马氏距离AUROC(ISIC测试集vs. PAD-UFES数据集)在所有水平上保持平稳且低于随机水平(约0.40),与所测试的五种几何指标均无显著关联。对余弦到质心得分器、合并k近邻得分器,以及三种非基于距离的得分器(基于能量的置信度得分、虚拟对数匹配、核密度估计器)也观察到相同的失效情况。8个得分器中有7个收敛到相同结果;基于能量的得分呈现孤立的上升趋势,我们报告该趋势但不将其视为反对整体模式的证据。无法访问训练目标的监督探针,在相同嵌入上以所有水平的0.72-0.81 AUROC恢复了领域成员身份,表明相关信息并未在表征中缺失。这些发现表明,仅分类性能可能会忽略学习到的表征中的信息是否以下游可靠性估计器可使用的形式组织。信息可以保持可解码,同时在很大程度上对非探针可靠性估计器不可访问。

英文摘要

Distance-based reliability estimation assumes that a representation's geometry reflects its trustworthiness, yet this assumption is rarely tested under training interventions that reshape geometry directly. We audit this assumption under domain-adversarial representation learning using a disentanglement dose-response ladder. Three checkpoint families share the same architecture and a 16-dimensional representation, differing only in orthogonality strength (lambda = 0, 1, 5). Representation geometry changed substantially with disentanglement strength: the condition number shifted by two orders of magnitude (Kendall tau = 0.84, exact p = 2.8e-5). This change was not accompanied by improved reliability estimation: Mahalanobis-distance AUROC (ISIC-test vs. PAD-UFES) remained flat and below chance (about 0.40) at every level, with no significant association with any of five geometry metrics tested. The same failure was observed for cosine-to-centroid and pooled k-nearest-neighbor scorers, plus three non-distance-based scorers: an energy-based confidence score, Virtual-Logit Matching, and a kernel density estimator. Seven of eight scorers converged on the same result; the energy-based score showed an isolated upward trend that we report but do not treat as evidence against the overall pattern. A supervised probe with no access to the training objective recovered domain membership from the identical embeddings at 0.72-0.81 AUROC across every level, showing that the relevant information was not absent from the representation. These findings indicate that classification performance alone can overlook whether information in a learned representation is organized in a form that downstream reliability estimators can use. Information can remain decodable while becoming largely inaccessible to non-probing reliability estimators.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑