自排斥采样用于扩散语言模型
Self-Repulsive Sampling for Diffusion Language Models
- ETH Zürich(苏黎世联邦理工学院)
- EPFL(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出自排斥采样器,通过同伴承诺在去噪步骤中惩罚重复标记,无需额外训练即可在扩散语言模型中生成多样化样本,显著提升投票准确率,在GSM8K上比贪心解码提高10个百分点以上。
中文摘要 AI 辅助
对多个响应进行采样并对其答案进行投票可以提高语言模型的准确性,但重复的答案限制了额外样本的收益。提高温度会增加多样性,但可能以牺牲单个样本的准确性为代价。我们引入了自排斥(SR),一种用于掩码扩散语言模型的采样器,它利用同伴承诺来使样本池多样化。在每个受惩罚的去噪步骤中,每条路径根据有多少同伴在同一位置承诺了该标记来降低该标记的对数几率。路径共享一个批处理前向传播,然后按顺序提交,因此后面的路径可以观察到同一步骤中较早做出的选择。这种耦合不需要训练或额外的前向或后向传播,即使在温度为零时也能产生不同的路径。当所有路径从相同的对数几率一起提交一个位置时,更新恰好最大化总对数几率减去一个凸重复成本。在LLaDA-8B-Instruct上,使用十条路径和128个去噪步骤,确定性SR在GSM8K上达到了80.38%的多数投票准确率,而未受惩罚的贪心解码器为70.17%。在温度为0.6且模型评估预算匹配的情况下,计数惩罚在32个块中比自一致性提高了2.06个百分点,在纯扩散下提高了14.50个百分点。在GSM8K、MATH和TruthfulQA上的实验表明,投票收益主要来自于正确答案覆盖率的提高,其收益因基准测试和解码方式而异。
英文摘要
Sampling several responses and voting over their answers can improve a language model's accuracy, but repeated answers limit the benefit of additional samples. Raising temperature increases diversity at a potential cost to per-sample accuracy. We introduce Self-Repulsion (SR), a sampler for masked diffusion language models that uses peer commitments to diversify the pool. At each penalized denoising step, each path lowers a token's logit according to how many peers have committed that token at the same position. Paths share a batched forward pass and then commit in sequence, so later paths observe choices made earlier in the same step. This coupling requires no training or additional forward or backward pass and can produce distinct paths even at temperature zero. When all paths commit a position together from identical logits, the update exactly maximizes total logit minus a convex duplication cost. On LLaDA-8B-Instruct with ten paths and 128 denoising steps, deterministic SR reaches 80.38% plurality accuracy on GSM8K, compared with 70.17% for the unpenalized greedy decoder. At temperature 0.6 and matched model-evaluation budgets, the count penalty improves over self-consistency by 2.06 percentage points in blocks of 32 and 14.50 under pure diffusion. Experiments on GSM8K, MATH and TruthfulQA show that voting gains arise mainly from higher coverage of correct answers, with gains that vary by benchmark and decoding regime.