AI 中文总结
本文针对度量修复问题,实现多类算法并评估其性能,发现修复效果主要取决于损坏类型与比例,且权重规则选择是关键,多数情况下修复未让下游任务结果更接近真实值。
AI 中文摘要
真实距离数据往往不满足要求:测量存在噪声、观测值缺失,最终得到的数值很少满足三角不等式。已有一系列方法用于修正这类数据,且所有这类方法都基于同一个假设——对修正后数据进行的分析,比基于原始数据的分析更能忠实反映真实情况。目前尚未有人验证过这一假设。我们通过度量修复问题来验证该假设,该问题要求找到最少的边,对其重新加权即可恢复三角不等式。我们实现了一套算法,涵盖现有文献中的方法及新提出的方法,包括带理论保证的算法和启发式算法,并在真实数据和合成数据(包括固有非度量数据和被损坏的数据)上评估其性能。我们发现,算法的性能主要由损坏的类型和比例决定,而非损坏程度或图的大小。我们进一步测试了修复对下游任务(即多维缩放(MDS)和k近邻(kNN))的影响,以判断修复后结果是否比损坏的实例更接近真实情况。在大多数情况下,修复并未使结果更接近真实值,我们确定了原因:仅靠少量边是不够的,找到正确的边集合(无论是注入的损坏还是自然的非度量性)至关重要。此外,权重规则的选择会影响性能:在有真实度量基准的实例中,度量修复算法可能会使图进一步偏离真实值,而获取真实权重的先验知识(oracle)会有所帮助,令人惊讶的是,相反情况也可能发生。设置权重并非实现细节,而是问题的一半。
英文摘要
Real distance data rarely cooperate: measurements are noisy, observations are missing, and the numbers that result seldom satisfy the triangle inequality. A family of methods exists to correct them, and every one of those methods rests on the same hope --- that analysis run on the corrected data is a more faithful surrogate for the truth than analysis run on the raw data. We are not aware of anyone having tested that hope. We test it through the problem of Metric Repair, which asks for the fewest edges whose reweighting restores the triangle inequality. We implement a suite of algorithms, covering the literature and new methods, both with theoretical guarantees and heuristics, and evaluate their performance on real and synthetic data, both inherently non metric and corrupted. We demonstrate that the algorithms' performance is determined predominantly by the type and fraction of corruption, rather than the corruption's magnitude or graph size. We further test the effect of repair on downstream tasks, namely MDS and $k$NN, and ask if the repair got the result closer to the truth compared to a corrupted instance. In most cases it did not, and we identify the culprit. A small set of edges is not enough. Finding the correct set of edges, be it an injected corruption or a natural non-metricity, is critical. Moreover, deciding on a weight rule impacts performance: on data instances with available metric ground truth, a metric repair algorithm can pull the graph further from the truth, while an oracle access to the true weights helps. Surprisingly, the opposite can be true as well. Setting the weights is not an implementation detail; it is half the problem.