黑板智能在全局约束问题上可超越自回归模型
Blackboard Intelligence Can Surpass Autoregressive on Globally Constrained Problems
浏览论文内容
中文总结 AI 辅助
针对全局约束问题,提出基于扩散语言模型的黑板智能推理方法,利用平均置信度指导搜索与修订,在多个基准上超越同规模自回归模型及部分前沿LLM。
中文摘要 AI 辅助
下一词预测推动了大型语言模型的显著进步,但越来越多的证据表明,它们在受复杂全局约束支配的问题上可能表现不佳。在这项工作中,我们聚焦于这一领域,并探究这些局限性是否部分源于下一词预测本身所引发的推理接口。我们通过黑板智能来研究这一问题:这是一种推理时视角,模型在一个固定且可修改的画布上工作,在候选解状态间进行搜索,而非承诺一条因果的、从左到右的轨迹。我们以扩散语言模型实例化这一思想,其任意顺序预测接口自然地暴露了对部分填充解状态的预测。我们的关键观察是,平均置信度——一个从标准掩码扩散目标中可获得的简单模型内部量——为全局连贯性提供了有用的代理,并能指导推理时的搜索与修订。实验上,在ZebraLogic、护士排班和作业车间调度上,黑板方法在保持微调后的LLaDA-8B-Instruct检查点不变的情况下持续改进推理,并大幅超越同规模自回归基线,在ZebraLogic-Hard上达到90.4%的准确率,在护士排班上达到76.4%的精确可行性,在JSSP上达到80.2%的最优性。更强的自回归搜索与细化也未能缩小在ZebraLogic-Hard上的差距,而黑板方法在ZebraLogic-Hard和JSSP上超越了所测试的前沿LLM,尽管这些模型规模大得多且具备强大的测试时推理能力。我们在以下URL开源了代码库:https://this https URL。
英文摘要
Next-token prediction has driven remarkable progress in large language models, yet a growing body of evidence suggests that they can struggle on problems governed by complex global constraints. In this work, we focus on this regime and ask whether some of these limitations arise from the inference interface induced by next-token prediction itself. We study this question through blackboard intelligence: an inference-time perspective in which a model works on a fixed, revisable canvas and searches over candidate solution states rather than committing to a causal, left-to-right trajectory. We instantiate this idea with diffusion language models, whose any-order prediction interface naturally exposes predictions over partially filled solution states. Our key observation is that mean confidence, a simple model-internal quantity available from the standard masked diffusion objective, provides a useful proxy for global coherence and can guide inference-time search and revision. Empirically, across ZebraLogic, Nurse Rostering, and Job-Shop Scheduling, Blackboard consistently improves inference while holding the fine-tuned LLaDA-8B-Instruct checkpoint fixed and substantially outperforms same-scale autoregressive baselines, reaching 90.4% accuracy on ZebraLogic-Hard, 76.4% exact feasibility on Nurse Rostering, and 80.2% optimality on JSSP. Stronger autoregressive search and refinement also fail to close the gap on ZebraLogic-Hard, while Blackboard surpasses tested frontier LLMs there and on JSSP despite their substantially greater scale and strong test-time reasoning. We open-source our codebase at https://github.com/jwoosang1/blackboard-intelligence.
发表机构
- Seoul National University(首尔大学)
- Harvard University(哈佛大学)
- University of Central Florida(中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。