发表机构
Thoughtworks(Thoughtworks)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究高温采样下大语言模型拒绝行为受温度影响的问题,提出顺序解码方法,能在高温时保留模型贪婪解码拒绝响应且延迟小,经实验在三个基准数据集上保留91 - 99%的拒绝行为,不影响安全提示高温响应。
AI 中文摘要
高温采样是增加大语言模型(LLMs)多样性的主要机制之一。基于截断的采样技术的最新进展有助于减轻高温采样的缺点,如神经文本退化,从而在不牺牲连贯性的情况下实现LLM输出的更多样性。然而,通过高温增加令牌概率分布的熵也会削弱模型的防护机制,降低其在有害提示下的拒绝响应。尽管高温采样有潜在好处且保持模型安全很重要,但缺乏在更高熵状态下维持LLMs拒绝行为的现有解决方案。为填补这一空白,我们系统研究温度如何影响LLMs中的拒绝行为,并提出一种有效的顺序解码方法,该方法在高温下保留模型的贪婪解码拒绝响应,同时增加的延迟最小。通过大量实验,我们表明我们的方法在三个基准数据集上保留了91 - 99%的贪婪解码拒绝行为,同时不影响模型对安全提示的高温响应。我们的工作展示了如何以有效方式为需要高温采样的应用维持拒绝行为。
英文摘要
Recent advances in truncation-based sampling have helped mitigate drawbacks of high-temperature sampling such as neural text degeneration, thereby enabling greater diversity without sacrificing coherence. However, increasing the entropy of the token probability distribution via high temperatures has also been shown to weaken the model's refusal response. Existing solutions for maintaining the refusal behavior of LLMs either replace the model's own refusal decision with a separate safety classifier or alter its output distribution for every prompt. To address this gap, we propose refusal-gated decoding (RGD): an efficient sequential decoding approach which preserves a model's greedy decoding refusal response at high temperatures and samples all other prompts from its exact direct high-temperature distribution, while incurring minimal additional latency. RGD runs a short greedy probe that reuses the prompt's KV cache and exits as soon as it becomes incompatible with a learned set of refusal prefixes; it returns the greedy response if the probe remains compatible and otherwise discards the probe and samples from the original prompt. Across seven models and three benchmark datasets at T=2.0, RGD raises greedy-refusal preservation from 91.9% under direct sampling to 98.3% on average while adding only 2.2-4.3% to the median per-request latency of non-refusals across temperatures. Unlike prompt-screening baselines which route many greedy non-refusals to greedy decoding, RGD keeps at least 98.1% of greedy non-refusals on unchanged high-temperature sampling, thereby preserving the model's natural high-temperature sampling behavior. We also propose a residual-stream variant of our method which lowers this latency overhead to at most 0.5% with comparable prompt routing accuracy. Our work shows that unlocking greater diversity via high-temperature sampling need not erode a model's refusal behavior.