arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

符号分支重复惩罚中的规范依赖性和结构化输出损坏:跨模型、推理堆栈和替代重复控制的测量

Gauge dependence and structured-output corruption in sign-branched repetition penalties: measurements across models, inference stacks, and alternative repetition controls

Peter Hollows

arXiv 2607.09791首次发表:更新:

AI 中文总结

研究符号分支重复惩罚中规范依赖性和结构化输出损坏问题,发现其惩罚定义不明确且损坏输出,将惩罚应用于归一化对数概率可消除这些问题,给出了相关机制、测量及归一化变体。

AI 中文摘要

乘法重复惩罚在大语言模型推理生态系统(如HuggingFace、vLLM等)中基于每个原始对数几率的符号进行分支(正数除以theta,负数乘以theta)。但由于softmax在对数几率上加常数不变,模型对数几率零点任意,符号分支读取该任意点。这导致惩罚定义不明确,例如在theta = 1时重新中心化对数几率无影响,但在theta = 1.3时会改变大量贪婪令牌;还会损坏结构化输出,如在200个真实世界JSON模式上,theta = 1.3会使有效且符合模式的输出率从97%降至23%。而将惩罚应用于归一化对数概率可消除这些影响,HuggingFace已提供该操作(LogitNormalization),目前默认关闭且在惩罚后应用。本文给出了机制、测量结果(在多个模型和数据集上)以及归一化变体。

英文摘要

The multiplicative repetition penalty shipped across the LLM inference ecosystem (HuggingFace, vLLM, llama$.$cpp, and a dozen further engines) branches on the sign of each raw logit (divide positives by theta, multiply negatives). But the softmax is unchanged by adding a constant to every logit, so a model's logit zero-point is arbitrary (a gauge choice), and the sign-branch reads it. Two measurable consequences follow. (1) The penalty is not well-defined: re-centering a model's logits by a constant is a provable no-op at theta=1, yet at a routine theta=1.3 it changes 58-96% of greedy tokens, while subtractive and normalized penalties change none; real checkpoints sit at widely different zero-points, so a fixed repetition_penalty is a different operation on every model. (2) It corrupts structured output: on 200 real-world JSON schemas, theta=1.3 drops the rate of valid, schema-conformant output from 97% to 23%. Applying the penalty to normalized log-probabilities instead of raw logits removes the gauge dependence by construction; HuggingFace's beam search has applied its processor chain, penalty included, to log-probabilities since at least v4.0.0, so repetition_penalty already names two different operators depending on decoding strategy. Because equal theta is not equal strength across the two operators, we also compare them at matched suppression, calibrated per model by search: there the normalized operator is statistically no worse on any quality metric measured, but on four of seven models it cannot match the raw operator's suppression at theta >= 1.15, and on six of seven at theta=1.3, the setting where the corruption was measured. This note gives the mechanism, the measurements (five models up to 7B; two code models; both effects replicated inside vLLM and llama$.$cpp through their own samplers), the per-model calibration map, and the normalized variant.

Comments10 pages, 2 figures, 6 tables. v2: adds matched-suppression comparison, per-model calibration map, and mechanism test. Code, data, per-stack survey and git genealogy: https://github.com/captainpete/repetition-penalty-gauge

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑