发表机构
Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统考察文本分类器在服务上下文变化下的非确定性,发现批形状和精度(如bf16)可显著影响预测分数与标签,并提出了保证标签稳定性的条件与缓解措施。
AI 中文摘要
确定性推理对于可靠和可信的机器学习至关重要。先前关于文本生成的研究表明,即使提示词、模型参数和采样随机性固定,改变批大小、批组成、硬件或推理引擎等因素也会改变生成的文本。这些差异部分归因于浮点非结合性、形状相关的内核选择以及数值执行中的其他实现级差异。然而,这些相同因素是否、何时以及在何种程度上影响文本分类仍不清楚。我们对文本分类器中的服务上下文非不变性进行了系统性研究,而先前的工作仅通过生成的文本对此进行测量。我们训练了180个模型,涵盖判别式、伪生成式和全生成式分类器形式,并在四类服务上下文中对每个模型进行评估,同时保持检查点和文本固定。标签稳定性并不暗示分数稳定性。在fp32比较中,仅改变批形状不会改变任何标签,但在bf16下,它移动了高达56.7个百分点的预测概率质量,标签变化集中在较小的余量上。在相同的服务变化下,全生成式分类器比判别式对应模型改变更多标签。我们推导了每种服务变化下标签稳定性的充分条件,并为每种机制提供了单独的缓解措施。我们的结果识别并量化了必须固定的服务条件,以实现可复现的文本分类。
英文摘要
Deterministic inference is essential for reliable and trustworthy machine learning. Prior studies of text generation have shown that changing factors such as batch size, batch composition, hardware, or inference engine can alter the generated text, even when the prompt, model parameters, and sampling randomness are fixed. These differences have been attributed in part to floating-point non-associativity, shape-dependent kernel selection, and other implementation-level differences in numerical execution. However, it remains unclear whether, when, and to what extent the same factors affect text classification. We present a systematic study of serving-context non-invariance in text classifiers, which prior work has measured only through generated text. We train 180 models spanning discriminative, pseudo-generative, and fully generative classifier formulations and evaluate each across four categories of serving contexts, holding the checkpoint and the text fixed. Label stability does not imply score stability. Changing only the batch shape changes no labels across fp32 comparisons, yet under bf16 it moves up to 56.7 percentage points of predicted probability mass, with label changes concentrated at small margins. Fully generative classifiers change more labels than their discriminative counterparts under the same serving changes. We derive sufficient conditions for label stability under each serving change and give a separate mitigation for each mechanism. Our results identify and quantify the serving conditions that must be fixed for reproducible text classification.