arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于掩码扩散语言模型的计算感知重掩码评估协议

CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models

Yash Shah, Abhijit Chakraborty, Vivek Gupta

arXiv 2607.24763首次发表:更新:

发表机构

Arizona State University; MongoDB(亚利桑那州立大学; MongoDB公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对掩码扩散语言模型评估标准不完善的问题,提出计算感知评估框架CaRE,通过标准化函数评估实际数量等方法审核重掩码策略,应用于多种模型和数据集,揭示了温度、计算匹配等因素对评估的影响,发布相关内容确保评估可重复可比。

AI 中文摘要

掩码扩散语言模型(MDLMs)发展迅速,但可靠解释其进展所需的评估标准却未能跟上。尽管MDLMs已能与自回归语言模型竞争,但近期七篇重掩码论文在不兼容的设置下进行评估,未联合控制名义步数、指标和采样温度等因素,致使策略排名难以比较。我们提出了CaRE,这是一个计算感知评估框架,通过标准化函数评估实际数量(NFE)、强制多指标报告以及明确控制随机性来审核MDLM重掩码策略。将其应用于LLaDA - 8B - Base和Dream - 7B - Base的7种重掩码策略,在OpenWebText和LM1B上的4个随机性水平和3个步长预算下,发现温度解释了大部分MAUVE方差;计算匹配比较逆转了一些已发表的策略排名;知情重掩码和随机解掩码存在冲突,高熵重掩码在256步、解掩码温度为0.25时使MAUVE降低0.296(p = 0.020)。涵盖12个开放权重MDLMs(参数从150M到8B)的CaRE排行榜表明这种交互方向在不同架构和规模中都成立。这些发现表明当前MDLM评估可能将算法改进与计算和随机性的隐藏选择系统地混为一谈。我们发布了评估协议、实现和排行榜,以确保未来重掩码声明具有可重复性和可比性。

英文摘要

Masked diffusion language models (MDLMs) are advancing rapidly, yet the evaluation standards needed to reliably interpret their progress have not kept pace. Despite MDLMs becoming competitive with autoregressive language models, seven recent remasking papers evaluate under incompatible settings, varying nominal step counts, metrics, and sampling temperatures without jointly controlling these factors, rendering their strategy rankings largely incomparable and leaving open whether reported gains reflect algorithmic improvements or evaluation artifacts. We present CaRE, a compute-aware evaluation framework that audits MDLM remasking strategies by standardizing actual number of function evaluations (NFE), enforcing multi-metric reporting, and explicitly controlling stochasticity. Applied to 7 remasking strategies across LLaDA-8B-Base and Dream-7B-Base at 4 stochasticity levels and 3 step budgets on OpenWebText and LM1B, CaRE reveals that: (i) temperature explains the majority of MAUVE variance, (ii) compute-matched comparisons reverse several published strategy rankings, and (iii) informed remasking and stochastic unmasking are in tension, with high-entropy remasking reducing MAUVE by 0.296 at 256 steps at unmask_temp=0.25 (p=0.020). A CaRE leaderboard covering 12 open-weight MDLMs (150M to 8B parameters) shows that this interaction direction holds across architectures and scales. These findings demonstrate that current MDLM evaluations can systematically conflate algorithmic improvements with hidden choices of compute and stochasticity. We release the evaluation protocol, implementation, and leaderboard to ensure future remasking claims are reproducible and comparable.

Commentsupdated version

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑