熵可以流动,也可以引导。成为熵。LEDFlow:将熵引导的生成顺序引入均匀离散流
Entropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete Flow
- University of California, Los Angeles(加州大学洛杉矶分校)
- Hong Kong University of Science and Technology(香港科技大学)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对均匀离散流中后续错误破坏正确中间预测的问题,提出LEDFlow采样器,利用局部熵自适应排序吸收顺序,在推理和生成任务上显著提升性能。
AI中文摘要:
均匀离散流允许在每个生成位置进行重复更新。虽然持续修正有助于纠正错误的标记,但也使正确的中间预测暴露于后续错误之中。在数独谜题上的一项实验表明,9.4%的生成单元在中间步骤是正确的,但在最终输出中却是不正确的。我们通过选择性吸收将生成顺序引入均匀离散流,该方法在固定所选预测的同时,保持活跃位置上的均匀流速度。为防止吸收错误的预测,我们提出了低熵离散流(LEDFlow),这是一种无需训练的采样器,通过局部熵自适应地排序吸收过程。通过将吸收误差分解为联合依赖项和条件预测项,我们证明了在固定吸收预算下选择最低熵位置能够最小化条件项的上界。我们进一步支持局部熵的选择,表明在非完美去噪器下,全局前瞻的决策误差界随前瞻窗口的增长而增大。在推理基准测试中,LEDFlow在Nikoli数独上达到了0.845的求解准确率,在强约束任务上取得了最大的提升。在文本到图像生成中,它取得了最佳总体得分,在多模态理解方面,它在全部六个基准测试上均优于原生采样器,且推理成本与标准流采样相当。
英文摘要:
Uniform discrete flow permits repeated updates at every generation position. While continued revision supports correction of wrong tokens, it also exposes correct intermediate predictions to later errors. An experiment on Sudoku puzzles shows that 9.4% of generated cells are correct at an intermediate step but incorrect in the final output. We introduce generation order into uniform discrete flow through selective absorption, which fixes chosen predictions while preserving the uniform flow velocity at active positions. To prioritize reliable predictions for absorption, we propose Low-Entropy Discrete Flow (LEDFlow), a training-free sampler that adaptively orders absorption by local entropy. By decomposing absorption error into joint dependence and conditional prediction terms, we show that, under entropy-error regularity, selecting the lowest-entropy positions under a fixed absorption count minimizes an upper bound on the conditional term. We further analyze sensitivity of global lookahead, whose worst-case decision-error bound grows with lookahead window under an imperfect denoiser. Across reasoning benchmarks, LEDFlow attains 0.845 Nikoli Sudoku solve accuracy, with largest gains on strongly constrained tasks. On a text-to-image generation benchmark it attains the best overall score among decode-time samplers, and on multimodal understanding it improves over the default sampler on all six benchmarks, at an inference cost comparable to standard flow sampling.