指令堆叠崩溃:基准测试集与提示词编译的能力依赖价值
Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation
AI总结:
本研究构建了含24个指令的基准测试集,发现指令遵循率随堆叠指令数增加非线性下降,提出无需训练的指令编译器可提升较弱生产级LLM的指令遵循率,且该收益与模型能力相关。
AI中文摘要:
生产环境中的提示词很少仅包含单一指令,一条系统消息可能同时要求输出有效的JSON、符合字数限制、包含3篇引用文献以及采用固定语气。本研究探讨当此类约束条件累积时,指令遵循能力会如何下降。我们引入了一个基准测试集,该测试集将24个经验证者检查的指令以1至20个为一组进行堆叠,并评估了三款生产级大语言模型(Claude Sonnet 4.6、GPT-5-mini、Gemini 2.5 Flash)。研究发现,指令遵循能力呈非线性下降:遵循率从约96%降至低至20%,其驱动因素是一组结构化且可复现的成对冲突;例如,单一的“输出JSON”约束与另外9个约束无法同时满足。随后,我们评估了一种无需训练的补救方法:指令编译器,它会在一次大语言模型调用中重写堆叠后的提示词,并可在多个查询中复用。该方法的收益具有能力分级特性:它为较弱模型恢复了高达11个百分点的遵循率,而这类模型也是最常被大规模部署的模型;对于已内化相同结构的较强模型,其遵循率基本不受影响。通过聚类稳健性测试、同基准对照组以及同系列缩放阶梯分析,我们将收益归因于重写本身,而非额外token、重排序或测量空间。我们发布了该基准测试集、验证器及缓存运行结果,以支持完整复现。
英文摘要:
Production prompts rarely carry a single instruction. One system message may require valid JSON, a word limit, three citations, and a fixed tone at the same time. We study how instruction-following degrades as such constraints accumulate. We introduce a benchmark that stacks 24 verifier-checked instructions, one to twenty at a time, and evaluate three production-tier LLMs (Claude Sonnet 4.6, GPT-5-mini, Gemini 2.5 Flash). Instruction-following degrades non-linearly: the follow rate falls from ~96% to as low as 20%, driven by a structured and reproducible set of pairwise conflicts. A single "output JSON" constraint, for example, is jointly unsatisfiable with nine others. We then evaluate a training-free remedy: an instruction compiler that rewrites the stacked prompt in a single LLM call and is reused across queries. Its benefit is capability-graded. It recovers up to +11 points of follow rate for weaker models, which are also the models most often deployed at scale, while leaving stronger models, which already internalise the same structure, essentially unchanged. Cluster-robust tests, same-baseline controls, and a within-family scaling ladder attribute the gain to the rewrite itself rather than to additional tokens, reordering, or measurement headroom. We release the benchmark, verifiers, and cached runs for full reproduction.