发表机构
University of Reading; Miralta Finance Bank S.A.; Albert School(雷丁大学; Miralta金融银行股份有限公司; 阿尔伯特商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对域偏移下后训练量化模型家族的选择问题,提出教师锚定方法,结合无标签失真与监督项,在134个候选家族上以最小标签预算降低平均遗憾。
AI 中文摘要
压缩已训练模型会产生一系列部署候选,而在域偏移下,最压缩的候选不一定是最适合部署的。我们研究此类候选家族的选择问题,其中候选和教师固定,目标标签缺失或稀缺。两个发现组织了无标签情况。最小教师失真几乎表现为恒定规则,在每次运行中都选择相同的八位、逐通道、未裁剪配置,但这并未最小化经验目标交叉熵。已有估计器明显分化:在CNN家族的过度自信崩溃机制中,基于置信度的估计器几乎将家族排序倒置,而识别该问题的诊断需要该设置所拒绝的标签,而输出分布估计器则匹配教师相对锚点,并在一种架构上优于它。尽管如此,失真是稳定的,因此监督项可以推动选择远离它。结合两者,我们给出了家族典型二次模拟的精确二次恒等式。我们还表明,在对称损坏下,线性于标签指示符的准则的标签相关部分乘以一个共同因子,只要其系数和是候选不变的,该类包含教师对比和准确性,但不包含交叉熵。这些刻画了得分的组成部分,而不限制选择遗憾。在一百三十四个候选家族中,每个家族对应一个独立训练的卷积或视觉Transformer教师,锚定在最小标签预算下减少了每个设置的平均遗憾,这一优势在超过二十五个标签后消失。
英文摘要
Compressing a trained model yields a family of deployment candidates, and under domain shift the most compressed one need not be the one to deploy. We study selection over such a family, with candidates and teacher fixed and target labels absent or scarce. Two findings organize the label-free case. Minimum teacher distortion behaves almost as a constant rule, selecting the same eight-bit, per-channel, unclipped configuration in every run, which does not minimize empirical target cross-entropy. Established estimators divide sharply: in the overconfident-collapse regime of the CNN families, confidence-based estimators order the family close to backwards, and the diagnostics that identify it need the labels the setting denies, while output-distribution estimators match the teacher-relative anchor and on one architecture beat it. Distortion is nonetheless stable, so a supervised term can move selection away from it. Combining the two, we give exact quadratic identities for a canonical quadratic analogue of the family. We also show that under symmetric corruption the label-dependent part of a criterion linear in the label indicator is multiplied by one common factor whenever its coefficient sums are candidate-invariant, a class holding teacher contrasts and accuracy but not cross-entropy. These characterize the score's components without bounding selection regret. Across one hundred and thirty-four candidate families, one per independently trained convolutional or Vision Transformer teacher, anchoring reduces mean regret at the smallest label budget in every setting, an advantage that fades beyond twenty-five labels.
Comments19 Pages, 3 Figures, 17 Tables