发表机构
Chongqing University of Posts and Telecommunications; Chongqing Normal University(重庆邮电大学; 重庆师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于粒球计算的自适应KNN方法,通过训练阶段构建多粒度粒球、预测阶段动态确定有效k值,在多数据集上的准确率和效率均优于现有KNN变体。
AI 中文摘要
k近邻(KNN)算法广泛应用于各类任务,k值的选择是关键问题,因其会显著影响性能。本文提出一种基于粒球计算的自适应高效KNN方法,该方法包含两个阶段:训练阶段,先对数据集进行粗划分以降低粒球内数据分布的复杂度,再引入Fisher准则控制粒球的分裂与停止,得到多粒度粒球表示;预测阶段,先通过加权距离机制定位最近粒球,再围绕测试样本构建自适应邻域,该邻域的实际样本数动态确定有效k值,最近粒球诱导的邻域能提供更稳定的局部组信息,从而提升对噪声和局部扰动的鲁棒性。实验结果表明,所提方法在多个数据集上的准确率和效率均优于现有KNN变体,代码已开源以支持可复现性:this https URL。
英文摘要
The $k$-Nearest Neighbor~(KNN) algorithm is widely used across various tasks. The selection of the $k$ value is a key issue because it significantly impacts performance. In this paper, an adaptive and efficient KNN approach via granular-ball computing is proposed. The method consists of two stages. \textcolor{black}{In the training stage, the dataset is first coarsely partitioned to reduce the complexity of data distributions within a granular ball, and then the Fisher criterion is introduced to control ball splitting and stopping, yielding a multi-granularity granular ball representation. In the prediction stage, the nearest granular ball is first located through a weighted distance mechanism, and an adaptive neighborhood is then constructed around the test sample. The effective $k$ value is dynamically determined by the actual number of samples contained in this neighborhood. The neighborhood induced by the nearest granular ball provides more stable local group information, thereby improving robustness against noise and local perturbations.} Experimental results demonstrate that the proposed method outperforms existing KNN variants across multiple datasets in terms of both accuracy and efficiency. The code has been open-sourced for reproducibility: https://github.com/lianxiaoyu724/Adaptive-GBKNN.