arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ProtoSeam:利用潜在高斯混合模型提升分类器训练

ProtoSeam: Lifting Classifier Training with Latent Gaussian Mixture Models

Robert Lampel, Timon Klein, Sebastian Sager

arXiv 2609.35174首次发表:更新:

发表机构

Otto von Guericke University; Max Planck Institute for Dynamics of Complex Technical Systems(奥托·冯·居里克大学; 马克斯·普朗克复杂技术系统动力学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ProtoSeam 提出一种提升式分类训练方法,通过在语义接口插入原型并采用一致性惩罚,在不改变推理架构下,将测试准确率提升最多五个百分点,并给出理论依据。

AI 中文摘要

我们提出了一种监督分类的提升式重构方法,在不改变推理时架构的情况下,提升标准分类器的最终准确率。网络 $N=N_2\circ N_1$ 在单一语义接口处被分割,并在该处为每个类别插入一个可学习的原型。训练结合了一个二次一致性惩罚项,将 $N_1(x)$ 拉向其类别的原型,以及一个在原型周围采样样本上评估的 $N_2$ 分类损失,其中梯度不穿过接口。推理时丢弃原型,使用未修改的网络 $N_2\circ N_1$。在 CIFAR-10、CIFAR-100 和 TinyImageNet 上,使用 ResNet 和视觉 Transformer 骨干,在共享调参协议下,提升式训练相比未提升的变体,测试准确率最高提升五个百分点。此外,我们为这些结果提供了理论证明。

英文摘要

We propose a lifted reformulation of supervised classification that improves the final accuracy of standard classifiers without changing the architecture at inference time. A network $N=N_2\circ N_1$ is split at a single semantic interface and one learnable prototype per class is inserted there. Training combines a quadratic consensus penalty that pulls $N_1(x)$ toward the prototype of its class with a classification loss of $N_2$ evaluated on samples drawn around the prototypes, whereat no gradient crosses the interface. At inference the prototypes are discarded and the unmodified network $N_2\circ N_1$ is used. Across CIFAR-10, CIFAR-100, and TinyImageNet with ResNet and vision transformer backbones, lifted training improves test accuracy by up to five percentage points over variants without lifting under a shared tuning protocol. Moreover, we provide theoretical justification of those results.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑