发表机构
Faculty of Information Technology, Monash University(莫纳什大学信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究加权 KNN 回归和软标签预测的数据 Shapley 问题,提出首个伪多项式时间精确算法、认证 FPTAS 等方法,给出复杂度分析和扩展,发布开源库及基准真值,在下游检测中体现精确性价值。
AI 中文摘要
数据 Shapley 是关于训练点价值的标准原则性答案,其 k 近邻(KNN)专门化是实际中部署的版本。已知无加权 KNN 和加权 KNN 分类的精确算法,但加权 KNN 回归和软标签预测一直存在难题。本文给出了加权 KNN 回归数据 Shapley 的首个伪多项式时间精确算法,还有连续权重和目标的认证 FPTAS,以及复杂度分析和加权软标签多类扩展。发布了开源库和首个精确加权回归数据 Shapley 基准真值。在下游错误标签检测中,精确值与蒙特卡洛数据 Shapley 统计等效,精确性的价值在于确定性、认证误差界限和审计估计器的精确参考。
英文摘要
Data Shapley can be computed exactly and efficiently for unweighted K-nearest-neighbor (KNN) models, the basis of the popular KNN-Shapley. For weighted KNN, however, efficient exact algorithms are known only for hard-label classification. In regression, the prediction is a weighted average whose normalization term changes with every subset. The best exact algorithm then takes O(N^K) time for N training points. We show that this normalization term is not an obstacle. A subset affects the prediction only through two numbers: the total weight of its K nearest neighbors and their total weighted target. Once the weights are discretized, subsets with equal total weight share their normalization term and can be summed as a group. This gives an exact algorithm that is near-linear in N and polynomial in K and the number of weight levels, for real targets under the squared loss. The same idea gives exact values for soft-label classifiers in the same time, and speeds up the earlier exact algorithm for hard labels by a factor of N. For continuous weights, we give a deterministic approximation with a certified error bound, without discretizing the weights. On the downside, we prove that some exact values take exponentially many bits to write down, so the dependence on weight precision cannot be removed in general. In experiments on real datasets with up to 100,000 training points and with 10 or 100 validation points, we compare the exact values with Monte Carlo estimates computed in the same time. Monte Carlo ranks the points less and less like the exact values as the training set grows. Two of its runs also disagree on a large share of the top 10% of points.
Comments77 pages, 4 figures. Version 3 is a substantial revision: exact values in near-linear time, continuous weights without a grid, limits of exact computation, and new experiments up to 100,000 training points