arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

极端分类:每类仅一个训练样本即可超越随机猜测

Extreme classification: beating chance with one training example from each class

Kevin Bleakley, Aaditya Ramdas

arXiv 2609.20897首次发表:更新:

发表机构

Inria; Laboratoire de Mathématiques d’Orsay; Stanford University(法国国家信息与自动化研究所; 奥赛数学实验室; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究极端分类:仅每类一个样本,通过构造随机核规则和确定性序规则,证明在多数分布对下可超越随机猜测,但不存在对所有分布对都有效的规则。

AI 中文摘要

我们研究一个极简分类问题:给定来自两个未知分布P和Q的独立带标签观测X~P和Z~Q,以及一个独立的目标Y以等概率从P或Q中抽取,当P≠Q时,能否对Y进行严格优于随机猜测的分类?一近邻规则对于具有不同均值且协方差矩阵为正定且相同的每一对多元高斯分布都能成功,但对于实直线上的光滑密度甚至可能表现严格差于随机猜测。我们构造了一个固定的随机核规则,其期望准确率恰好为1/2+MMD_k^2(P,Q)/4,并从可数族可测二元问题中获得可数生成可测空间上的特征核。我们还证明了在实数集上,一个确定性序规则对每一对不同的Borel概率测度都能超越随机猜测。通过可测编码,我们得到一个确定性的无分布规则,该规则在每一个可数生成的可测空间(特别是每一个可分度量空间)上都能超越随机猜测。最后,我们证明没有任何规则能对每一对不同的分布和每一个未知的不平衡类别先验都有效;在自适应目标类别选择下,除了公平硬币之外的每个规则对某些有限支撑的分布对都严格差于随机猜测。

英文摘要

We study a minimal classification problem: Given independent labeled observations $X\sim P$ and $Z\sim Q$ from two unknown distributions $P,Q$, and given an independent target $Y$ drawn with equal probability from $P$ or $Q$, can one classify $Y$ strictly better than chance whenever $P\neq Q$? The one-nearest-neighbor rule succeeds for every pair of multivariate Gaussian distributions with distinct means and a common positive-definite covariance matrix but can perform strictly worse than chance even for smooth densities on the real line. We construct a fixed randomized kernel rule whose expected accuracy is exactly $1/2+\operatorname{MMD}_k^2(P,Q)/4$, and obtain characteristic kernels on countably generated measurable spaces from countable families of measurable binary questions. We also prove that a deterministic order rule on $\mathbb R$ beats chance for every pair of distinct Borel probability measures. A measurable encoding then gives a deterministic distribution-free rule which beats chance on every countably generated measurable space, in particular every separable metric space. Finally, we show that no rule works for every distinct pair of distributions and every unknown unbalanced class prior; under adaptive target-class selection, every rule other than a fair coin is strictly worse than chance for some finitely supported pair.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑