arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

XTC:通过排除顶部选择的头感知采样

XTC: Head-Aware Sampling by Excluding Top Choices

Philipp Emanuel Weidmann, Allen Roush, Judah Goldfeder, Sanjay Basu, Ravid Shwartz-Ziv

arXiv 2608.22758首次发表:更新:

发表机构

Thoughtworks; Columbia University; Oracle; New York University(Thoughtworks公司; 哥伦比亚大学; 甲骨文公司; 纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

XTC是一种轻量头感知解码算子,通过排除部分高概率词元改善自回归语言模型的多样性-重复帕累托前沿,在多项实验中提升创意生成效果且准确率损失极小,已被多个工具采用。

AI 中文摘要

自回归语言模型的标准解码规则通过重新缩放完整的下一个词元分布或截断其低概率尾部来促进多样性。这些策略忽略了一种常见的开放式生成场景:其中有多个合理的延续内容,但概率质量仍过多集中在最通用的选择上。我们引入XTC(Exclude Top Choices,排除顶部选择),一种轻量的头感知解码算子,直接针对该场景。XTC识别概率超过绝对合理性阈值τ的词元:当至少有两个符合条件时,它以概率ρ移除占主导的合格选择,仅保留最弱的合理替代项后再进行重新归一化。在针对Gemma 3 27B Q4、Gemma 3 12B Q6和DeepSeek R1 14B Q6开展的60项实验中,结合对Llama 3.3 70B Q4的缩放验证,XTC改善了多样性-重复的帕累托前沿。在创意生成任务上,四款模型的Distinct-2提升了11%至15%,重复三元组降低了27%至47%;与温度缩放结合后,相较于基线,Distinct-2的增益达38%,重复三元组降低率达71%。由150名Master评分者参与的盲测Amazon Mechanical Turk研究显示,XTC获得62.3%的创意偏好(p<10⁻⁴)且未降低流畅性,GPT-4o对照评判者在各项指标上均重现了Anthropic评判者的方向。在使用Llama 3.3 70B Q4的IFEval任务中,XTC将提示级严格准确率保持在基线的1.7个百分点以内,同时恢复了大部分多样性增益;而在Distinct-2上匹配的温度设置则使IFEval准确率降低了8.8个百分点。该效果与温度和重复惩罚呈叠加关系,在量化级别和模型族间均具有鲁棒性,且在12种提示类型中表现一致。XTC已被this http URL、ExLlamaV2和text-generation-webui采用。

英文摘要

Standard decoding rules for autoregressive language models promote diversity by rescaling the full next-token distribution or truncating its low-probability tail. These strategies overlook a common regime of open-ended generation in which several continuations are plausible but too much probability mass remains concentrated on the most generic choice. We introduce XTC (Exclude Top Choices), a lightweight head-aware decoding operator that targets this regime directly. XTC identifies tokens whose probabilities exceed an absolute plausibility threshold $τ$: when at least two qualify, it removes the dominant eligible choices with probability $ρ$ and retains only the weakest plausible alternative before renormalization. Across 60 experiments on Gemma 3 27B Q4, Gemma 3 12B Q6, and DeepSeek R1 14B Q6, with scaling validation on Llama 3.3 70B Q4, XTC improves the diversity-repetition Pareto frontier. On creative generation, Distinct-2 increases by 11--15% and repeat trigrams decrease by 27--47% across the four models. Combined with temperature scaling, gains reach 38% in Distinct-2 and 71% in repeat-trigram reduction over baseline. A blinded Amazon Mechanical Turk study with 150 Master raters yields a 62.3% creativity preference for XTC ($p<10^{-4}$) without reduced fluency, while a GPT-4o control judge reproduces the Anthropic-judge direction on every measure. On IFEval with Llama 3.3 70B Q4, XTC preserves prompt-level strict accuracy within 1.7 percentage points of baseline while recovering most of the diversity gain; a temperature setting matched on Distinct-2 reduces IFEval by 8.8 points. The effect is additive with temperature and repetition penalties, robust across quantization levels and model families, and consistent across twelve prompt genres. XTC has been adopted by llama.cpp, ExLlamaV2, and text-generation-webui.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑