arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不要重复你自己:在采样时停止逐字循环

Don't Repeat Yourself: Stopping Verbatim Loops at Sampling Time

Philipp Emanuel Weidmann, Allen Roush, Judah Goldfeder, Sanjay Basu, Ravid Shwartz-Ziv

arXiv 2608.22761首次发表:更新:

发表机构

Thoughtworks; Columbia University; Oracle; New York University(Thoughtworks; 哥伦比亚大学; 甲骨文公司; 纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型的逐字循环问题,提出采样时的DRY对数几率调整方法,在多模型和研究中有效降低循环率并保持性能,已被开源LLM推理框架采用。

AI 中文摘要

大语言模型以自回归方式生成文本,但开放式生成易出现逐字循环,即模型重复上下文中已有的文本片段。标准防御措施如重复惩罚、存在惩罚、频率惩罚和n元组阻塞针对的是词元重复而非循环的序列结构,通常只有在足够强的抑制强度下才能减少循环,但这会同时降低文本格式或流畅度。我们提出Don't Repeat Yourself(DRY),一种采样时的对数几率调整方法,仅在候选词元会扩展当前后缀为上下文早期出现的片段的精确延续时才对其进行惩罚。序列保护器可保护聊天模板和格式词元。在参数规模从15亿到1200亿的模型、9种提示族及600对人类研究中,DRY将后缀扩展率降低了47%,同时提升了词汇多样性。干预匹配安慰剂未产生可比的降低,表明后缀匹配是其作用机制。在经AWQ量化的700亿和1200亿参数模型上,DRY将循环率降低约一半,同时保持MT-Bench、MMLU和GSM8K的性能,而标准替代方法则出现可测量的性能下降。DRY已被包括this http URL、ExLlamaV2和text-generation-webui在内的流行开源LLM推理框架采用,凸显其对文本生成的实际影响。

英文摘要

Large Language Models generate text autoregressively, but open-ended generation is prone to verbatim looping, in which models repeat spans already present in context. Standard defenses such as repetition, presence, and frequency penalties and n-gram blocking act on token recurrence rather than the sequential structure of a loop, and often suppress looping only at strengths that also degrade formatting or fluency. We propose Don't Repeat Yourself (DRY), a sampling-time logit adjustment that penalizes a candidate token only when generating it would extend the current suffix into an exact continuation of a span seen earlier in the context. Sequence breakers protect chat templates and formatting tokens. Across models from 1.5B to 120B parameters, nine prompt families, and a 600-pair human study, DRY reduces suffix-extension rate by 47% while improving lexical diversity. An intervention-matched placebo produces no comparable reduction, identifying suffix matching as the operative mechanism. On AWQ-quantized 70B and 120B models, DRY reduces loop rate by roughly half while preserving MT-Bench, MMLU, and GSM8K performance, whereas standard alternatives lose measurable ground. DRY has been adopted by popular open-source LLM inference frameworks including llama.cpp, ExLlamaV2, and text-generation-webui, highlighting its practical impact on text generation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑