发表机构
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对冻结语言模型中行为规范与执行状态解耦问题,提出重放门控神经执行框架,通过状态索引认证集与预算受限查找分离规范与实现,实验验证其可靠性与成本优势。
AI 中文摘要
输入条件化的神经干预引发了一个运行时问题:当一种行为规范允许多个其有效性取决于执行状态的动作时,什么能够持久存在?我们引入了重放门控神经执行,它分离了五个对象:一个持久行为谓词、其状态索引的已认证实现集、一个瞬态动作见证、一个预算受限的查找器,以及执行授权。候选对象在冻结模型上进行隔离的FP32/BF16重放;提交承诺还需要有效的运行审计。在Qwen3-0.6B和SmolLM2-360M-Instruct上的实验确立了这些对象的不同失败模式。独立初始化在所有24个测试的固定状态单元中产生不同的已认证动作。未改变的SmolLM2见证在所有128个原生状态中保持已认证,但在384个非对角线转移中仅有66个通过。所有767个存档的Qwen见证均成功重放,但预算受限的查找器在三次预设运行中均遗漏了一个已知可实现的单元。在1,141个重放提交的候选中,174个未通过项目认证。一个冻结的三层级联使用这些边界来拒绝未认证的提案,并升级审计有效的搜索遗漏。在256个先前封存的Qwen Fresh请求中,221个首先在最低成本层级获得认证,所有256个均获得审计授权,未观察到绕过。相对于冻结的全量搜索,中位单例搜索与认证成本比为0.1055,P95为1.3485,包括失败的层级。在所研究的两个小模型的行为族内,这些结果支持状态索引的、集合值的执行语义:规范持久存在,搜索提出见证,重放认证加运行审计授予执行权限。
英文摘要
Input-conditioned neural interventions raise a runtime question: what persists when one behavioral specification admits multiple actions whose validity depends on execution state? We introduce replay-gated neural execution, separating five objects: a persistent behavioral predicate, its state-indexed certified realization set, a transient action witness, a budget-limited finder, and execution authorization. Candidates undergo isolated FP32/BF16 replay of the frozen model; commitment additionally requires a valid run audit. Experiments on Qwen3-0.6B and SmolLM2-360M-Instruct establish distinct failure modes for these objects. Independent initializations yield distinct certified actions in all 24 tested fixed-state cells. Unchanged SmolLM2 witnesses remain certified in all 128 native states but only 66 of 384 off-diagonal transfers. All 767 archived Qwen witnesses replay successfully, yet a budget-limited finder misses one known-realizable cell in all three prespecified runs. Of 1,141 replay-submitted candidates, 174 fail item certification. A frozen three-tier cascade uses these boundaries to reject uncertified proposals and escalate audit-valid search misses. On 256 previously sealed Qwen Fresh requests, 221 first certify at the lowest-cost tier and all 256 receive audited authorization, with no observed bypass. Relative to frozen full search, the median singleton search-and-certification cost ratio is 0.1055 and P95 is 1.3485, including failed tiers. Within the studied behavioral family on two small models, these results support state-indexed, set-valued execution semantics: specifications persist, search proposes witnesses, and replay certification plus run audit grants execution authority.