arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于降维保真度的固定半径距离带基准

A Fixed-Radius Distance-Band Benchmark for Dimensionality-Reduction Fidelity

Yoshio Takaeda

arXiv 2608.21779首次发表:更新:

发表机构

toor Inc.(toor公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出固定半径距离带Shepard ρ基准,对比8种DR方法,发现现有排名指标有偏,固定半径指标可揭示聚类压缩等问题,还能解决单点及少数总体的评估难题。

AI 中文摘要

降维(DR)方法通常通过每个点的k近邻在二维嵌入中的保留情况来评估(recall@k、可信度、连续性)。我们认为这类指标是距离保真度的有偏衡量:其逐点可变半径和硬包含阈值偏向邻接图方法(t-SNE、UMAP),并惩罚保留绝对距离的方法。我们改用固定半径距离带的Shepard ρ来评估DR保真度:高维与二维成对距离的斯皮尔曼相关性,限制在累积距离带内,以便分别报告近距和全局结构,且每个点都以相同的绝对半径进行评估。在具有已知真实几何结构的合成数据集(非均匀密度、密集聚类、闭环过渡、子空间外异常值、不平衡双总体数据)和现实噪声(SNR=1,D=768,N=1000)下,我们对8种方法进行基准测试——PCA、Isomap、t-SNE、UMAP、PyMDE、PCC、DREAMS及闭源的toorPIA,结果显示:(i)高全局Shepard ρ可与聚类内规模约93倍的压缩共存,该压缩在基于排名的指标中不可见,但在基于值的过度压缩指标中明显;(ii)recall@k与固定半径带存在系统性分歧,分歧方向与偏差预测一致;(iii)受成员限制的Shepard ρ可解决多对统计无法处理的单点和少数总体问题,而近期的局部-全局混合方法DREAMS在这些问题上表现失效。补充的样本外(addplot)测试用于判断从未见过的异常是否落在正常区域外,以及其方向是否能识别来源。所有指标均基于所有成对距离精确计算,与任何方法的内部结构无关,且所有数值可离线复现:闭源方法的输出坐标(而非其算法)已提交至工件。

英文摘要

Dimensionality-reduction (DR) methods are routinely judged by how well each point's k nearest neighbors survive the 2-D embedding (recall@k, trustworthiness, continuity). We argue this family is a biased measure of distance fidelity: its per-point variable radius and hard inclusion threshold favor neighbor-graph methods (t-SNE, UMAP) and penalize methods that preserve absolute distances. We instead score DR fidelity with a fixed-radius distance-band Shepard rho: the Spearman correlation between high-D and 2-D pairwise distances, restricted to cumulative distance bands so that near and global structure are reported separately, with every point judged on the same absolute radius. On synthetic datasets with known ground-truth geometry (non-uniform density, dense clusters, a closed-loop transition, off-subspace outliers, imbalanced two-population data) at realistic noise (SNR=1, D=768, N=1000), we benchmark eight methods -- PCA, Isomap, t-SNE, UMAP, PyMDE, PCC, DREAMS, and the closed-source toorPIA -- and show that (i) high global Shepard rho can coexist with a ~93x collapse of within-cluster scale, invisible to rank-based scores but obvious in a value-based over-compression metric; (ii) recall@k and the fixed-radius band disagree systematically, in the direction the bias predicts; (iii) a membership-restricted Shepard rho resolves single-point and minority-population questions that many-pair statistics cannot -- questions on which even DREAMS, a recent local-plus-global hybrid, fails silently. A supplementary out-of-sample (addplot) test asks whether a never-seen anomaly lands outside the normal region and whether its direction identifies its source. All metrics are computed exactly on all pairwise distances, independently of any method's internals, and every number is reproducible offline: the closed-source method's output coordinates (not its algorithm) are committed to the artifact.

Comments20 pages, 13 figures. Code and data: https://doi.org/10.5281/zenodo.21380823

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑