相同探针,不同数字:激活探针对推理时数值非确定性的鲁棒性如何?
Same Probe, Different Numbers: Are Activation Probes Robust to Inference-Time Numerical Non-Determinism?
浏览论文内容
中文总结 AI 辅助
本研究评估激活探针在推理时数值非确定性下的鲁棒性,发现探针稳定但聚合准确率低估判定变化,建议用逐示例一致性评估并区分表征噪声与输入变化。
中文摘要 AI 辅助
激活探针越来越多地被用于监控部署中的大语言模型(LLM)。探针通常在一种推理配置下训练,然后在服务栈所使用的任意批大小和数值精度下使用。由于常见的GPU内核不具备批不变性,且浮点格式的舍入方式不同,部署时看到的激活值与探针训练时的激活值并不相同。我们测量了这在Llama-3.1-8B、Qwen3-8B和Gemma-3-4B上,在批大小4、8和16以及float32、bfloat16和float16下所造成的代价,在一种配置上训练了768个探针,在每种其他配置上评估每个探针,并逐示例比较判定结果。探针是稳定的,但聚合准确率是展示这种稳定性的错误工具:它低估了判定变化的数量,实际变化因子为2到9倍。在提示阶段,跨1,392次转移,准确率变化从未超过0.47个百分点,仅0.076%的判定发生变化;在float32下仅改变批大小时,201,960个判定中无一变化。在解码过程中,翻转率升至2.8%,但实际生成的token与参考匹配的行中,翻转仅占0.12-0.15%的情况,而token分化的行中翻转占12.9%:原因是文本,而非算术。bfloat16批大小变化导致2.1%的行的首个生成token翻转,到第20个token时,25%的行处于不同的token上。翻转是对称的,Cohen's kappa保持在0.94以上,AUROC变化最多0.05点。在底层,激活的变化幅度与格式的舍入误差相当:bfloat16批大小变化对激活的中位相对L2扰动为1e-2,约为float16数值的8倍。探针吸收了这些变化;而模型自身的下一个token的argmax则不然。激活监控器的鲁棒性评估应报告逐示例的一致性,而非聚合准确率,将表征噪声与输入变化分开,并说明服务配置。
英文摘要
Activation probes are increasingly used to monitor LLMs in deployment. A probe is typically trained under one inference configuration, then used under whatever batch size and numerical precision the serving stack uses. Because common GPU kernels are not batch-invariant and floating-point formats round differently, the activations seen at deployment are not the ones the probe was trained on. We measure what that costs for Llama-3.1-8B, Qwen3-8B and Gemma-3-4B across batch sizes 4, 8 and 16 and float32, bfloat16 and float16, training 768 probes on one configuration, evaluating each on every other, and comparing verdicts example by example. Probes are stable, but aggregate accuracy is the wrong instrument for showing it: it understates how many verdicts change by a factor of two to nine. At the prompt, accuracy never moves by more than 0.47 percentage points across 1,392 transfers and only 0.076% of verdicts change; under float32 with only the batch size varied, none of 201,960 verdicts change. During decoding the flip rate rises to 2.8%, but rows whose realised tokens matched flip in only 0.12-0.15% of cases, while rows whose tokens diverged flip in 12.9%: the cause is the text, not the arithmetic. A bfloat16 batch-size change flips the first generated token for 2.1% of rows and leaves 25% on different tokens by token 20. Flips are symmetric, Cohen's kappa stays above 0.94, and AUROC moves by at most 0.05 points. Underneath, activations move about as much as the format's rounding: a bfloat16 batch-size change perturbs them by a median relative L2 of 1e-2, roughly 8x the float16 figure. Probes absorb this; the model's own next-token argmax does not. Robustness evaluations of activation monitors should report per-example agreement rather than aggregate accuracy, separate representational noise from input change, and state the serving configuration.