AI 中文总结
针对扩散大语言模型去噪调度器存在的EOS溢出和近端偏差问题,提出一种基于CMA-ES的轻量级进化启发式调度器,仅需393个参数,在多个基准上超越现有方法。
AI 中文摘要
扩散大语言模型(dLLMs)最近作为传统自回归(AR)大语言模型(LLMs)的一种有前景的替代方案而出现。通过利用双向注意力和并行解码,dLLMs实现了更高效的生成。然而,它们在推理时(训练时不存在)需要一个精心设计的去噪调度器,其选择对生成质量有显著影响。虽然基于置信度的启发式调度器已显示出强大的经验性能,但它们存在两个关键失败模式:EOS溢出和近端偏差。通过对Transformer注意力模式的深入分析,我们揭示这些失败源于某些位置对无效标记(如[MASK]和[EOS])分配了不成比例的高注意力权重,从而产生误导性的置信度信号。基于这一见解,经验证据表明,有效的注意力分数可以为传统的基于置信度的启发式方法提供补充指导,但没有任何单一指标在所有场景中始终表现出色,这意味着最优去噪轨迹高度依赖于上下文。为解决此问题,我们提出了一种使用协方差矩阵自适应进化策略(CMA-ES)优化的轻量级进化启发式调度器。我们的调度器动态整合多个启发式特征与上下文平均场嵌入,仅需393个可训练参数。在LLaDA和Dream上跨四个推理和规划基准的评估中,我们的方法始终优于强基线,包括传统启发式方法、块自回归方法和近期最先进(SOTA)方法。据我们所知,它代表了迄今为止参数效率最高的神经调度器。我们的代码可在以下网址获取:https URL。
英文摘要
Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to conventional Auto-Regressive (AR) Large Language Models (LLMs). By leveraging bidirectional attention and parallel decoding, dLLMs enable more efficient generation. However, they require a carefully designed denoising scheduler at inference time (absent during training) whose choice significantly impacts generation quality. While confidence-based heuristic schedulers have shown strong empirical performance, they suffer from two critical failure modes: EOS Overflow and Proximal Bias. Through in-depth analysis of the Transformer's attention patterns, we reveal that these failures stem from certain positions assigning disproportionately high attention weights to invalid tokens (e.g., [MASK] and [EOS]), which produce misleading confidence signals. Building on this insight, empirical evidence shows that valid attention scores can provide complementary guidance to conventional confidence-based heuristics, yet no single metric consistently excels across all scenarios, implying that the optimal denoising trajectory is highly context-dependent. To address this problem, we propose a lightweight evolutionary heuristic scheduler optimized using the Covariance Matrix Adaptation Evolution Strategy (CMA-ES). Our scheduler dynamically integrates multiple heuristic features with a contextual mean-field embedding, while requiring only 393 trainable parameters. Evaluated on LLaDA and Dream across four reasoning and planning benchmarks, our method consistently outperforms strong baselines, including conventional heuristics, block auto-regressive methods, and recent State-Of-The-Art (SOTA) approaches. To the best of our knowledge, it represents the most parameter-efficient neural scheduler to date. Our code is available at https://github.com/RS2002/Evo-Denoise .