AI 中文总结
提出ScopeIF框架,通过统一约束分解和分级奖励建模,提升大语言模型在复杂范围感知指令上的精确遵循能力,实验显示优化后的小模型可媲美前沿大模型。
AI 中文摘要
精确指令遵循是大语言模型(LLMs)的一项基本能力,要求其输出严格满足输入指令中的客观约束。在复杂应用场景中,这些约束通常具有多样的范围,它们支配特定的响应片段而非整个输出。然而,现有的优化方法在数据构建过程中往往忽视约束范围,并依赖二值化的逐约束奖励,这导致数据多样性有限且对复杂约束的监督稀疏。为此,我们提出了ScopeIF,一种用于范围感知的精确指令遵循的新型训练框架。我们首先引入一个统一模式,将客观约束分解为三个解耦维度:范围(Scope)、目标(Target)和区间(Range)。基于该模式,我们构建了ScopeInstruct,一个包含多样化范围感知约束的大规模指令数据集,并将基于工具的验证与分级奖励建模相结合,以量化每个约束的违反程度,为策略优化提供密集监督。大量实验表明,ScopeIF在性能上持续优于现有方法,尤其是在复杂的范围感知约束上,同时保持了一般能力。值得注意的是,它使优化后的Qwen3-4B和8B模型能够与Gemini-2.5-Pro和DeepSeek-V3.2等强大的前沿模型相媲美或超越,为推进范围感知指令遵循建立了一种有效范式。我们的代码和数据可在该https URL获取。
英文摘要
Precise instruction-following is a fundamental ability of large language models (LLMs), requiring their outputs to strictly satisfy objective constraints in input instructions. In complex application scenarios, these constraints often possess diverse scopes that govern specific response segments rather than the entire output. However, existing optimization methods often neglect constraint scope during data construction and rely on binary per-constraint rewards, yielding limited data diversity and sparse supervision for complex constraints. To this end, we propose ScopeIF, a novel training framework for scope-aware precise instruction-following. We first introduce a unified schema that factorizes objective constraints into three decoupled dimensions: Scope, Target, and Range. Grounded in this schema, we construct ScopeInstruct, a large-scale instruction dataset with diverse scope-aware constraints, and combine tool-grounded verification with graded reward modeling to quantify the violation degree of each constraint, providing dense supervision for policy optimization. Extensive experiments demonstrate that ScopeIF consistently outperforms existing methods, particularly on complex scope-aware constraints, while preserving general capabilities. Notably, it enables optimized Qwen3-4B and 8B models to rival or surpass strong frontier models such as Gemini-2.5-Pro and DeepSeek-V3.2, establishing an effective paradigm for advancing scope-aware instruction-following. Our code and data are available at https://github.com/thu-coai/ScopeIF.
Comments28 pages, 8 figures