arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于自我改进系统的可证伪发布门限

Falsifiable Release Gates for Self-Improving Systems: Standing Invariants at Scale

Deepak Soni

arXiv 2607.13070首次发表:更新:

发表机构

AI Architect(人工智能架构师)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究自我改进系统安全声明,提出可证伪发布门限及构建验证方法,在Antahkarana运行时应用,经机器检查确保安全,发布验收结果,明确范围,使结果可重现、门限可在其他框架运行。

AI 中文摘要

关于自我改进智能体运行时的安全声明几乎总是自我分级的,如策略文件、护栏或自述承诺。我们描述了可证伪发布门限以及构建和验证此类系统的方法,即每个新功能在发布前必须通过预先指定的、机器可验证的验收套件,且在每个门限处保留一组固定的不变量。我们在Antahkarana开放运行时中应用了该方法,通过七个门限从基本可观测性进入自我管理循环,该循环可对自身策略提出更改建议。对有界模型的一百万个记录可达状态空间进行详尽的机器检查,确保没有安全关键属性能力令牌的操作不会发送到执行器,并根据执行跟踪重新检查。发布了所有七个门限的验收测量结果,明确了每个声明的范围,使结果可重现,门限可在其他智能体框架上运行。

英文摘要

Safety claims for self-improving agent runtimes are almost always self-graded: a policy file, a guardrail, a promise in a README. We describe falsifiable release gates, a methodology in which every new capability must pass a pre-declared, machine-checkable acceptance suite before it ships, while a fixed set of standing invariants is preserved across every gate. We instantiate it in Antahkarana, an open runtime, then do what a method paper is only vindicated by: we follow the same runtime as it grows and ask whether the guarantees survive. The safety-critical property, that no action reaches an effector without a capability token minted by a control ring, is machine-checked exhaustively over the reachable states of a bounded model; a deliberately broken model yields the shortest counterexample, so the checker demonstrably has teeth. We then carry the runtime through six further releases. Across every one, the action-safety invariants INV-1 through INV-6 held without a single change, and one release added three capabilities while introducing no new invariant. Under the same teeth discipline, six more machine-checked families were added: memory with provable unlearning, a governed agent, calibrated abstention over a post-quantum record, a harness of many sub-agents, the self-improvement loop itself, and the residency of what it produces. The acceptance suite grew from 122 tests to 563. The load-bearing result sits in the negative space: across more than a doubling of capability, the safety core was neither weakened nor redesigned. The last families are the first on real hardware: gated self-improvement compounds a small model from 20% to 70% accuracy while auto-rejecting a candidate that only inflates confidence, and the whole governed path costs 0.021 ms per request, 0.008% of model inference. We release the runtime, tools, and gate suite; every number reproduces with a single command.

Comments23 pages, 14 figures. Major revision merging the follow-up "Standing Invariants at Scale" into this paper per arXiv moderation: the machine-checked action-safety core preserved across six further releases, six new invariant families with teeth, and real-hardware self-improvement; suite grown from 122 to 563 tests. Software, gate suite, run artifacts, and TLA+ spec are open source

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑