发表机构
ADA University(阿塞拜疆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种基于精确求解子样本的集成方法用于聚类回归,通过投票或选择组合结果,无需预设修剪水平,在高达20%异常值时准确率达0.89,优于传统修剪交替法。
AI 中文摘要
聚类最小二乘法将回归数据划分为K组,每组具有独立的线性拟合。我们研究了一种集成方法,其基学习器是精确的:在B个大小为m << n的随机子样本上,将问题全局最优求解,通过最近曲面分配扩展每个解,对齐标签,并通过投票或选择组合重复结果。每个重复结果随后成为m个点上的经验K-量化器,且该集成允许精确分析。一旦干净子样本的概率乘以干净数据重复结果的准确率超过二分之一,投票在某个单元上即为正确,无论污染值如何;在集成未标记的单元上迭代该集成,可得到一种估计修剪水平而非要求修剪水平的变体。当响应中存在高达20%的严重异常值时,其最差情况准确率为0.89,而真实污染比例下的修剪交替法为0.80,且在所有尝试的固定水平下均更低。在数据条件下,重复结果是独立同分布的,因此投票在B上以指数速度收敛到重复结果分布的多数分区,与边界集之外的标准最小化器一致,该边界集的大小取决于m和微求解器,而非B。对于两组位置模型,阶数为1/pi_min的子样本在分离阈值以上使得重复结果正确的频率高于错误的频率。O(n^3)枚举给出了两组和一个协变量的精确最小化器,理论与之对照验证。在干净数据上,该集成在相同成本下输给多起点交替法。
英文摘要
Trimming methods for robust clusterwise regression discard a fixed fraction of the data. Too low a level breaks the fit; too generous a level can trim away a small group. We first propose a group-recovery step that can follow any trimming or flagging method: it searches the discarded units for a line, tests whether those near it form a peak rather than a band, and restores the line as a group when the likelihood of Gaussian groups plus uniform noise improves by a margin like that of the Bayesian information criterion. In simulations it repaired the failures of a generous TCLUST-REG level with unequal groups, but not with three or four groups. The second proposal, ESF (exact-subsample flagging), is a flagging procedure without a trimming level: it solves the clusterwise least-squares problem exactly on small subsamples, flags units far from the best fit, and draws later subsamples from the rest. Two constants stand in for the level: a subsample size, set from a lower bound on the smallest group proportion, and a cap on the flagged set. It is meant for data about whose contamination nothing is known: TCLUST-REG at the fixed level 0.30, followed by reweighting and the recovery step, was as accurate as ESF on average up to a fifth of outliers, and higher levels were more accurate beyond. The flagged fraction estimates the contamination only when the errors are close to Gaussian. On taxi fares with known tariffs, ESF found both in every sample.
Comments27 pages, 2 figures; supplementary material (54 pages) as an ancillary file. Version 2 is a rewritten paper with a new title. Code and raw results: https://github.com/sorujov/eesp, doi:10.5281/zenodo.23046086