arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02639cs.SEcs.AI

指令堆叠崩溃:基准测试集与提示词编译的能力依赖价值

Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation

Atul Anand, Sourav Chattaraj

AI总结:

本研究构建了含24个指令的基准测试集,发现指令遵循率随堆叠指令数增加非线性下降,提出无需训练的指令编译器可提升较弱生产级LLM的指令遵循率,且该收益与模型能力相关。

AI中文摘要:

生产环境中的提示词很少仅包含单一指令,一条系统消息可能同时要求输出有效的JSON、符合字数限制、包含3篇引用文献以及采用固定语气。本研究探讨当此类约束条件累积时,指令遵循能力会如何下降。我们引入了一个基准测试集,该测试集将24个经验证者检查的指令以1至20个为一组进行堆叠,并评估了三款生产级大语言模型(Claude Sonnet 4.6、GPT-5-mini、Gemini 2.5 Flash)。研究发现,指令遵循能力呈非线性下降:遵循率从约96%降至低至20%,其驱动因素是一组结构化且可复现的成对冲突;例如,单一的“输出JSON”约束与另外9个约束无法同时满足。随后,我们评估了一种无需训练的补救方法:指令编译器,它会在一次大语言模型调用中重写堆叠后的提示词,并可在多个查询中复用。该方法的收益具有能力分级特性:它为较弱模型恢复了高达11个百分点的遵循率,而这类模型也是最常被大规模部署的模型;对于已内化相同结构的较强模型,其遵循率基本不受影响。通过聚类稳健性测试、同基准对照组以及同系列缩放阶梯分析,我们将收益归因于重写本身,而非额外token、重排序或测量空间。我们发布了该基准测试集、验证器及缓存运行结果,以支持完整复现。

英文摘要:

Production prompts rarely carry a single instruction. One system message may require valid JSON, a word limit, three citations, and a fixed tone at the same time. We study how instruction-following degrades as such constraints accumulate. We introduce a benchmark that stacks 24 verifier-checked instructions, one to twenty at a time, and evaluate three production-tier LLMs (Claude Sonnet 4.6, GPT-5-mini, Gemini 2.5 Flash). Instruction-following degrades non-linearly: the follow rate falls from ~96% to as low as 20%, driven by a structured and reproducible set of pairwise conflicts. A single "output JSON" constraint, for example, is jointly unsatisfiable with nine others. We then evaluate a training-free remedy: an instruction compiler that rewrites the stacked prompt in a single LLM call and is reused across queries. Its benefit is capability-graded. It recovers up to +11 points of follow rate for weaker models, which are also the models most often deployed at scale, while leaving stronger models, which already internalise the same structure, essentially unchanged. Cluster-robust tests, same-baseline controls, and a within-family scaling ladder attribute the gain to the rewrite itself rather than to additional tokens, reordering, or measurement headroom. We release the benchmark, verifiers, and cached runs for full reproduction.

补充信息

↑