扩散语言模型中的原位指令跟随
In-Place Instruction Following in Diffusion Language Models
浏览论文内容
中文总结 AI 辅助
提出原位指令跟随任务及IIF-Bench基准,并设计GRAFT后训练框架,通过约束感知SFT和偏好优化,在四个扩散语言模型上将IIF分数平均提升15.35分。
中文摘要 AI 辅助
扩散大语言模型(dLLMs)通过双向迭代去噪生成文本,天然支持锚定在任意输出位置的用户指定约束,这种范式被称为原位提示(IPP)。我们将其形式化为原位指令跟随(IIF)任务,并构建了IIF-Bench,一个涵盖字面、风格和话语功能约束的分层基准,并配以基于评分标准的局部-全局评估协议。一种推理时的注意力偏置探针表明,普通dLLMs在去噪过程中往往对约束跨度关注不足。随后,我们提出了GRAFT,一个面向IPP的后训练框架,结合了约束感知的SFT和偏好优化。在四个代表性的dLLMs上,GRAFT将平均IIF分数从57.75提升至73.10(+15.35分),在字面和话语功能约束上分别获得15.91和15.57分的绝对提升,同时保持了一般生成能力。
英文摘要
Diffusion Large Language Models (dLLMs) generate text via bidirectional iterative denoising, naturally supporting user-specified constraints anchored at arbitrary output positions, a paradigm known as In-place Prompting (IPP). We formalize this as the In-place Instruction Following (IIF) task and construct IIF-Bench, a hierarchical benchmark spanning literal, style, and discourse-function constraints, paired with a rubric-based local-global evaluation protocol. An inference-time attention-bias probe suggests that vanilla dLLMs often under-prioritize constraint spans during denoising. We then propose GRAFT, an IPP-oriented post-training framework combining constraint-aware SFT and preference optimization. On four representative dLLMs, GRAFT raises the average IIF score from 57.75 to 73.10 (+15.35 points), with absolute gains of 15.91 and 15.57 points on literal and discourse-function constraints, while preserving general generation ability.
发表机构
- National University of Singapore(新加坡国立大学)
- Nanyang Technological University(南洋理工大学)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。