发表机构
University of Cambridge; Maccabim-Re’ut High School; The Open University of Israel(剑桥大学; 马卡比姆-鲁特高中; 以色列开放大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM防护的自适应系统存在的裁决陈旧性问题,提出新鲜度受限防护(FBS)方法,可显著降低批准过期率,同时制定新鲜度契约规范批准的有效性。
AI 中文摘要
用于自适应系统(SAS)的大语言模型(LLM)防护机制可能在检查时发出正确的批准,但在执行时已失效,这会产生执行阶段的检查时间到使用时间(TOCTOU)风险。我们研究裁决新鲜度,即防护裁决在使用时是否仍然有效。我们区分三个回答不同问题的指标:固定动作重放下的全候选裁决变化率、记录的闭环轨迹上的神谕标记批准过期率,以及基于判断者条件的使用时无效率。在五个可复现的SAS环境中,当重放偏移量为8个模拟器步长时,全候选裁决变化率在5.3%-48.4%之间。我们提出新鲜度受限防护(FBS),它从批准的安全侧裕度和近期特征波动估计每个批准的有效时间范围,无需显式的植物动力学模型。使用工件中记录的固定设置,FBS在相同偏移量下将神谕标记批准过期率从3.4%-24.7%降至0-1.8%。对四个LLM判断者的单独审计发现,每个批准流中都存在非零的基于判断者条件的使用时无效率。我们制定了新鲜度契约:每个批准必须在检查时正确且在使用时保持有效。
英文摘要
A large language model (LLM) guardrail for a self-adaptive system (SAS) may issue an approval that is correct at check time but stale by actuation. This creates an Execute-stage time-of-check to time-of-use (TOCTOU) hazard. We study verdict freshness: whether a guardrail verdict remains valid when used. We distinguish three quantities that answer different questions: all-candidate verdict change under fixed-action replay, oracle-labeled approval expiry on recorded closed-loop trajectories, and judge-conditioned use-time invalidity. Across five reproducible SAS environments, all-candidate verdict-change rates span 5.3-48.4% at a common replay shift of eight simulator steps. We introduce the Freshness-Bounded Shield (FBS), which estimates each approval's validity horizon from its safe-side margin and recent feature volatility, without an explicit plant-dynamics model. Using fixed settings documented in the artifact, FBS reduces oracle-labeled approval-expiry rates from 3.4-24.7% to 0-1.8% at the same shift. A separate audit of four LLM judges finds nonzero judge-conditioned use-time invalidity in every approval stream. We formulate a freshness contract: every approval must be correct at check time and remain valid at use time.
CommentsAccepted at the 2026 IEEE International Conference on Autonomic Computing and Self-Organizing Systems (ACSOS 2026)