arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从接口到推理:从任意阶模型中引出任意阶推理

From Interface to Inference: Eliciting Any-Order Inference from Any-Order Models

Seunggeun Kim, Jaeyeon Kim, Taekyun Lee, Yuyuan Chen, Yilun Du, Sham Kakade, Sitan Chen

arXiv 2607.26504首次发表:更新:

发表机构

University of Texas at Austin; Harvard University(德克萨斯大学奥斯汀分校; 哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对离散推理任务的任意阶推理需求,提出两种掩码扩散改进方法,训练对应模型验证了方法可诱导任意阶推理行为并提升下游性能。

AI 中文摘要

代码生成等许多离散推理任务本质上是非因果的:程序员在高级结构和局部细节之间切换,这一过程我们称为任意阶推理。对于缺乏原生任意阶接口的自回归语言模型,补全和下一个编辑预测等非因果能力需要手动设计的机制。我们能否转而设计原生支持任意阶推理的模型?掩码扩散模型近期成为极具吸引力的候选,因为其任意阶训练目标自然提供了任意阶预测接口。然而,该接口不会自动产生任意阶推理。我们证明这一接口-推理差距源于位置不确定性:固定画布的 token 级模型可能知道应出现何种语义组件,却不知道将其放置在何处。据此,我们提出两种互补方法:(1)基于插入的掩码扩散,以 FlexMDM(Kim 等人,2025)为基础,通过插入放宽固定位置约束,支持跨非连续区域生成;(2)潜空间掩码扩散,将预测转移到更粗粒度的语义片段,支持在潜生成顺序上搜索。实验上,我们训练了一个用于 Python 编码的 7B FlexMDM 和一个用于 GSM8K 的 125M LatentMDM,结果显示两种方法均诱导出不同的任意阶推理行为,并提升了下游性能。我们在该 https URL 发布代码库。

英文摘要

Many discrete reasoning tasks, such as code generation, are inherently non-causal: programmers move between high-level structure and local details, a process we call any-order inference. For autoregressive language models, which lack a native any-order interface, non-causal abilities such as infilling and next-edit prediction require hand-designed mechanisms. Can we instead design models that natively support any-order inference? Masked diffusion models have recently emerged as compelling candidates, as their any-order training objective naturally offers an any-order prediction interface. This interface, however, does not automatically yield any-order inference. We demonstrate that this interface-inference gap stems from positional uncertainty: fixed-canvas, token-level models may know what semantic component should appear without knowing where to place it. In light of this, we propose two complementary approaches: (1) Insertion-based masked diffusion, building on FlexMDM (Kim et al, 2025), relaxes fixed-position commitments via insertions, enabling generation across non-contiguous regions. (2) Latent-space masked diffusion shifts prediction to coarser semantic segments, enabling search over latent generation orders. Empirically, we train a 7B FlexMDM for Python coding and a 125M LatentMDM for GSM8K and show that both approaches induce distinct any-order inference behaviors and improve downstream performance. We release our codebase at https://github.com/SeunggeunKimkr/genuine-any-order.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑