用集成方法可视化非线性降维中的不确定性
Visualizing Uncertainty in Non-linear Projections with Ensembles
浏览论文内容
中文总结 AI 辅助
该研究针对UMAP、t-SNE等非线性降维方法的投影变异性问题,提出用多个NLDR输出的中位数可视化、输入扰动生成共识嵌入的方法,可更好传达投影模式可靠性。
中文摘要 AI 辅助
UMAP和t-SNE等广泛使用的非线性降维(NLDR)方法具有随机性,对同一数据重复运行会产生不同的低维投影。本文探讨与投影变异性相关的两个问题:部分数据集的聚类、结构和离群点在不同运行间可能发生变化,而另一些数据集的投影在过拟合噪声时会极其稳定。针对第一个问题,我们提出可视化多个NLDR输出的中位数,而非依赖单个投影;针对第二个问题,我们在创建共识嵌入前对输入数据进行扰动。我们发现,多个投影的中位数在多个质量指标上的表现与单个运行相当,同时增加扰动会增强全局结构而非局部结构。通过一系列探索性可视化,我们表明即使是相对简单的集成呈现也能更好地传达投影模式的可靠性。
英文摘要
Widely used non-linear dimensionality reduction (NLDR) methods such as UMAP and t-SNE are stochastic--repeated runs on the same data can produce different low-dimensional projections. In this paper, we explore two problems related to projection variability: on some datasets clusters, structure, and outliers may change run-to-run, and on others projections can be extremely stable when overfitting noise. To address the first problem, we propose visualizing the median of multiple NLDR outputs rather than relying on individual projections. To address the second, we perturb input data before creating consensus embeddings. We find that taking the median of multiple projections performs comparably to individual runs on multiple quality metrics, while increasing perturbation emphasizes global over local structure. We show through a set of exploratory visualizations that even relatively simple ensemble presentations can be used to better communicate the reliability of projection patterns.