发表机构
University of Sheffield(谢菲尔德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对黑盒大语言模型对话中撤销约束失效的问题,提出 sysname 方法,通过合约账本、顺序消融探针和修复阶梯实现复发率的测量、预测与修复,降低了约束撤销的失效概率。
AI 中文摘要
多轮对话中,用户施加约束和撤销约束的操作同样便捷,但撤销操作并不能可靠生效:模型仍会执行已撤回的要求(偶尔会在声明已移除的评论之下),这种失效被称为“行为复发”或“撤销惯性”。目前尚无工具可测量每个条款的影响、在交付前预测该影响,或在匹配预算下修复该问题。\n\nsysname 仅通过模型 API 填补了这三个空白:合约账本将每个约束与可执行检查器配对,将撤销操作记录为“墓碑”,并提前将净约束状态编译为单个规范;顺序消融探针测量每个条款的依从性和增量行为效应;修复阶梯在匹配的 token 数量和尝试次数预算下运行。\n\n在 dataname(NTasks 个人评估任务,NClauses 已验证检查器)上,当约束负载从 ScaleDelayedMTwo 增长到 ScaleDelayedMEight 时,8B 操作点的复发率上升,而更强的模型则处于基准水平。在匹配的检查器、模型和预算下,与无账本验证器重试基线相比,提前编译显著降低了复发率(RestoreDiff,95% 置信区间 RestoreDiffCI,p 值 RestoreDiffP);在此基础上叠加的自适应阶梯干预未产生可检测的增益(95% 置信区间排除了增益≥ LadderExcludedGain 的情况)。该探针可在交付前预测复发率(AUROC AurocPrimary);单句墓碑注释可恢复约三分之一的编译效果,且能通过安慰剂对照。在 CostDeliveryFactor 的交付开销和每个结果的 API 计算 CostTotalHedged 下,撤销失效成为对话状态可测量、可预测和可修复的属性,而非不可见属性。
英文摘要
Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enacting withdrawn requirements (occasionally beneath comments asserting their removal), a failure we call \emph{behavioral relapse}, or revocation inertia. No existing instrument measures this influence per clause, predicts it before delivery, or repairs it under matched budgets. \sysname{} closes the three gaps through the model API alone: a contract ledger pairs every constraint with an executable checker, records revocations as tombstones, and compiles the net constraint state ahead of time into a single specification; a sequential ablation probe measures per-clause adherence and incremental behavioral effect; a repair ladder operates under token- and attempt-matched budgets. On \dataname{} (\NTasks{} HumanEval tasks, \NClauses{} verified checkers), relapse at an 8B operating point climbs from \ScaleDelayedMTwo{} to \ScaleDelayedMEight{} as constraint load grows, while stronger models sit at floor. Under matched checkers, model, and budget, ahead-of-time compilation significantly reduces relapse against a no-ledger verifier-retry baseline (\RestoreDiff{}, 95\% CI \RestoreDiffCI{}, $p$ \RestoreDiffP{}); adaptive ladder interventions stacked on top add no detectable gain (95\% confidence excludes gains $\geq$ \LadderExcludedGain{}). The probe predicts relapse before delivery (AUROC \AurocPrimary{}); a one-sentence tombstone note recovers about a third of the compilation effect and survives a placebo control. At \CostDeliveryFactor{} delivery overhead and \CostTotalHedged{} of API compute for every result, revocation failure becomes a measurable, predictable, and repairable property of dialogue state rather than an invisible one.