arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07160cs.CLcs.AI

扩散语言模型中的原位指令跟随

In-Place Instruction Following in Diffusion Language Models

Zheng Nie, Zherui Li, Jiaming Zhang, Kun Wang, Zhenhong Zhou, Yufei Guo

首次发表
浏览论文内容

中文总结 AI 辅助

提出原位指令跟随任务及IIF-Bench基准,并设计GRAFT后训练框架,通过约束感知SFT和偏好优化,在四个扩散语言模型上将IIF分数平均提升15.35分。

中文摘要 AI 辅助

扩散大语言模型(dLLMs)通过双向迭代去噪生成文本,天然支持锚定在任意输出位置的用户指定约束,这种范式被称为原位提示(IPP)。我们将其形式化为原位指令跟随(IIF)任务,并构建了IIF-Bench,一个涵盖字面、风格和话语功能约束的分层基准,并配以基于评分标准的局部-全局评估协议。一种推理时的注意力偏置探针表明,普通dLLMs在去噪过程中往往对约束跨度关注不足。随后,我们提出了GRAFT,一个面向IPP的后训练框架,结合了约束感知的SFT和偏好优化。在四个代表性的dLLMs上,GRAFT将平均IIF分数从57.75提升至73.10(+15.35分),在字面和话语功能约束上分别获得15.91和15.57分的绝对提升,同时保持了一般生成能力。

英文摘要

Diffusion Large Language Models (dLLMs) generate text via bidirectional iterative denoising, naturally supporting user-specified constraints anchored at arbitrary output positions, a paradigm known as In-place Prompting (IPP). We formalize this as the In-place Instruction Following (IIF) task and construct IIF-Bench, a hierarchical benchmark spanning literal, style, and discourse-function constraints, paired with a rubric-based local-global evaluation protocol. An inference-time attention-bias probe suggests that vanilla dLLMs often under-prioritize constraint spans during denoising. We then propose GRAFT, an IPP-oriented post-training framework combining constraint-aware SFT and preference optimization. On four representative dLLMs, GRAFT raises the average IIF score from 57.75 to 73.10 (+15.35 points), with absolute gains of 15.91 and 15.57 points on literal and discourse-function constraints, while preserving general generation ability.

发表机构

  • National University of Singapore(新加坡国立大学)
  • Nanyang Technological University(南洋理工大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑