超越块边界:扩散大语言模型的多块编辑
Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models
浏览论文内容
中文总结 AI 辅助
研究针对块扩散语言模型的块边界问题,提出多块编辑(MBE)方法。通过无需训练的解码算法编辑前序块令牌,引入监督微调策略,扩展SGLang提升效率。实验表明该方法在保持吞吐量时优于基线,在长程一致性任务上有显著性能提升。
中文摘要 AI 辅助
块扩散已成为扩展离散扩散语言模型(dLLMs)的主导范式,因其在固定大小块中解码文本可保持块内并行生成且二次注意力成本可控。然而,这一效率伴随着结构限制,即块末尾的令牌生成时无法访问未来的跨块上下文,且块一旦确定,其不确定预测对后续块就不可逆转。这导致了块边界问题,不确定性在块边界累积,早期错误在后续生成中传播。为解决此问题,我们提出多块编辑(MBE),通过基于跨块上下文编辑解码令牌来缓解该问题。MBE首先提出一种无需训练的解码算法来编辑前序块中的解码令牌,通过在选定块上重新打开全注意力窗口实现。鉴于块扩散训练与MBE推理间注意力机制不匹配,MBE进一步引入监督微调策略,为模型配备双向注意力掩码以逐步扩大编辑跨度。此外,它还通过多形状CUDA图池和细粒度KV缓存控制扩展SGLang,使这些可变长度编辑过程在实践中高效。在13个基准上对LLaDA2.1-Mini进行的实验表明,无需训练的MBE在保持可比吞吐量的同时优于所有现有解码基线,MBE SFT进一步带来2.7的性能提升。最大的改进出现在需要强长程一致性的任务上,如在AIME 2025上提升13.3,在ZebraLogic上提升5.9,证明了MBE的有效性。
英文摘要
Block diffusion is the dominant approach for scaling discrete diffusion language models (dLLMs), as fixed-size blocks preserve parallel decoding while keeping quadratic attention costs tractable. Yet blockwise generation creates a structural weakness: tokens near a block boundary lack future cross-block context, and errors in finalized blocks become irreversible context for later generation. We call this the block boundary problem. Measuring predictions with and without later-block context shows that boundary sensitivity rises sharply: on AIME 2025, mean self-containedness divergence (SCD) in the last quarter of a block is 61.3 times that in the first quarter. We propose Multi-Block Editing (MBE), which revises decoded tokens using cross-block context. Training-Free MBE reopens a full-attention window over selected blocks without parameter updates. To address the mismatch between block-diffusion training and MBE inference, Multi-Block Edit SFT introduces bidirectional attention masks and progressively enlarges the editing span. We also extend SGLang with a multi-shape CUDA Graph pool and fine-grained KV-cache control for efficient variable-length editing. Experiments on LLaDA2.1-Mini across 12 benchmarks show broad, consistent gains. Training-Free MBE improves or matches standard decoding on every benchmark. Full MBE raises the 12-benchmark average from 61.45 to 64.24, with gains of up to 13.33 points on AIME 2025, while retaining 87.3--96.7% of standard-decoding end-to-end throughput across four datasets.
发表机构
- Ant Group(蚂蚁集团)
- Fudan University(复旦大学)
- Shanghai Innovation Institute(上海创新院)
机构由 AI 辅助整理,请以论文原文为准。