arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

贝叶斯熵重排序用于校准扩散语言模型

Bayesian Entropy-based Reordering for Calibrated Diffusion Language Models

Zhejun Jiang, Mijung Park

arXiv 2610.05125首次发表:更新:

发表机构

The University of British Columbia(不列颠哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出BayesER,一种基于贝叶斯后验熵的扩散语言模型解码框架,通过不确定性引导标记提交,降低校准误差并保持或提升准确性,且不确定性信号可跨数据集迁移。

AI 中文摘要

掩码扩散语言模型(MDLMs)通过迭代地将掩码标记替换为模型预测来生成序列。在每个去噪步骤中,解码器选择哪些位置具有足够的置信度以进行提交。现有的解码方法通常依赖于softmax置信度,这可能会被错误校准。我们引入了BayesER(贝叶斯熵重排序),一种事后贝叶斯解码框架,利用预测不确定性来指导标记提交。在BayesER中,我们构建了一个以预训练检查点为中心的轻量级近似后验,类似于Laplace-LoRA,但无需训练LoRA适配器。我们对后验样本的预测进行平均,并使用预测熵来优先考虑可靠的位置。我们研究了后验预测如何影响跨代码生成、数学推理、规划和分子生成基准的位置排序和标记选择。我们表明,与常见的解码方案(包括置信度阈值解码)相比,BayesER降低了序列级校准误差,同时保持或提高了准确性。此外,在一个代码生成数据集上拟合的后验无需重新拟合即可减少另一个数据集的校准误差,这表明贝叶斯不确定性可能为更可靠的MDLM解码提供可转移的信号。

英文摘要

Masked Diffusion Language Models (MDLMs) generate sequences by iteratively replacing masked tokens with model predictions. At each denoising step, the decoder chooses which positions are sufficiently confident to commit. Existing decoding methods typically rely on softmax confidence, which can be miscalibrated. We introduce BayesER (BAYESian Entropy-based Reordering), a post-hoc Bayesian decoding framework that uses predictive uncertainty to guide token commitment. In BayesER, we construct a lightweight approximate posterior centered at the pretrained checkpoint, similar to Laplace-LoRA but without training LoRA adapters. We average predictions over posterior samples and use predictive entropy to prioritize reliable positions. We examine how posterior predictions affect position ordering and token selection across benchmarks spanning code generation, mathematical reasoning, planning, and molecular generation. We show that BayesER reduces sequence-level calibration error while preserving or improving accuracy relative to common decoding schemes, including confidence-threshold decoding. Additionally, a posterior fitted on one code-generation dataset reduces calibration error on another without refitting, suggesting that Bayesian uncertainty may provide a transferable signal for more reliable MDLM decoding.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑