发表机构
Wageningen University & Research(瓦赫宁根大学与研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FloDR是基于可逆归一化流的降维方法,保留非嵌入坐标并生成可精确计算的诊断场,解决了t-SNE、UMAP等方法丢失支撑属性含义信息的问题。
AI 中文摘要
高维数据的二维嵌入常被解读超出其所能支撑的范围,t-SNE、UMAP等方法的输出中,簇内与簇间的距离、空白区域的含义、每个点隐藏的结构量通常不可见,这是因为优化过程中丢弃了能支撑这些属性含义的信息。本文提出FloDR,一种通过可逆归一化流嵌入数据的降维方法,FloDR仅用前两个输出坐标生成二维嵌入,保留其余坐标而非丢弃。训练得到的映射具备精确逆映射与精确密度属性,可从绘制布局的模型精确逆映射而非近似逆映射计算诊断可视化。具体而言,绘制两个场:条件扩散,以输入单位度量每个嵌入位置处原始数据的未确定量;隐藏对比度,度量绘制的两个坐标丢弃的标记对比度信息量。两个场均针对保留的输入数据部分进行预设测试并附带自助法置信度,未通过测试的场会被标记为弃权(不执行)。
英文摘要
It is common for two-dimensional embeddings of high-dimensional data to be read far beyond what they can support. Distances in and between clusters, the meaning behind empty spaces, and the amount of structure hidden at each point are generally invisible in the output of methods such as t-SNE and UMAP. This is because the information that could support the meaning of these properties is discarded during the optimisation process. Here, we present FloDR, a dimensionality reduction method that embeds data through an invertible normalising flow. While FloDR only uses the first two output coordinates to create a two-dimensional embedding, it retains the remaining coordinates rather than discarding them. In addition to the embedding, an exact inverse and an exact density are properties of a trained mapping, which enable diagnostic visualisations that are computed from the exact inverse of the model that drew the layout rather than from an approximate one. Specifically, we draw two fields, the conditional spread, which measures how much of the original data remains undetermined at each embedding position in input units, and the hidden contrast, which measures how much information about a labelled contrast the two plotted coordinates discard. Both fields are rendered with a prespecified test against a held out portion of the input data and a bootstrap confidence. A field that fails the test is reported as refused.
Comments22 pages, 12 figures