arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12426cs.AIcs.CL

大语言模型能遵循指令,但无法同时遵循太多:组合约束满足中的相变

Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction

Mariya I. Vasileva

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出约束饱和评估基准,发现大语言模型在同时遵循5-6个以上指令约束时可靠性崩溃,不同类型约束的性能下降幅度存在差异,失败呈近独立的乘法累积效应。

中文摘要 AI 辅助

大语言模型正越来越多地被部署到需要同时遵守多个显式约束的场景中,这些约束包括推理结构、安全边界、输出模式。单个约束能被熟练处理,但多个约束需同时满足的组合场景仍缺乏充分刻画:性能会以多快的速度下降?下降由什么决定?能否缓解这种崩溃?我们引入了约束饱和评估(Constraint Saturation Evaluation, CSE),这是一个通过程序生成的基准,系统地改变同时约束的数量(k),每个约束由确定性的基于规则的验证器评分,不涉及任何大语言模型评判:共15个模型、36种约束类型,在k=1-12时进行了369753次检查。得出三个发现:第一,每个约束的通过率逐渐且可预测地下降,而满足所有k个约束的概率会崩溃——一个在k=8时单个约束通过率约为41%的模型,成功满足全部8个约束的概率仅为5.7%。第二,约束的下降程度并不相同:结构约束每增加一个约束,基线能力的损失是词汇约束的2倍,这一差异由理解-维持差距决定,该差距将需要持续跟踪的约束与不受组合影响的二元决策约束区分开。第三,失败几乎是独立的,这正是累积效应呈乘法性的原因;仅存的残余耦合与共享输出特征相关,而非成对干扰——错误的句子计数会导致所有读取它的约束失败。可靠的指令遵循在超过5-6个同时约束时会失效:最强模型在7个约束时的探测级成功率降至50%以下,而15个模型中有12个在3个或更少约束时就已降至该水平。

英文摘要

Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled proficiently, but the compositional regime, where many must hold jointly, remains poorly characterized: how rapidly does performance degrade, what governs the degradation, and can the collapse be mitigated? We introduce Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that systematically varies the number of simultaneous constraints (k), with every constraint scored by a deterministic, rule-based verifier and zero LLM-judge involvement: 15 models, 36 constraint types, 369,753 checks at k=1-12. Three findings emerge. First, per-constraint pass rate decays gradually and predictably, while the chance of satisfying all k constraints collapses - a model passing individual constraints at ~41% at k=8 succeeds on all eight just 5.7% of the time. Second, constraints do not degrade equally: structural constraints lose 2x more baseline capability per added constraint than lexical ones, ordered by a comprehension-maintenance gap that separates constraints requiring sustained tracking from binary decisions immune to composition. Third, failures are nearly independent, which is what makes the accumulation multiplicative; the residual coupling that does exist tracks shared output features rather than pairwise interference - a wrong sentence count fails every constraint that reads it. Reliable instruction following breaks down beyond 5-6 simultaneous constraints: probe-level success falls below 50% at 7 constraints for the strongest model, and at 3 or fewer for 12 of 15.

发表机构

  • Meta Superintelligence Labs(元超级智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑