Falsifiable Release Gates for Self-Improving Systems: Standing Invariants at Scale
用于自我改进系统的可证伪发布门限
机构 * AI Architect(人工智能架构师)
专题命中 安全训练 :safety(abstract,comments);分类 cs.AI
AI总结 研究自我改进系统安全声明,提出可证伪发布门限及构建验证方法,在Antahkarana运行时应用,经机器检查确保安全,发布验收结果,明确范围,使结果可重现、门限可在其他框架运行。
Comments 23 pages, 14 figures. Major revision merging the follow-up "Standing Invariants at Scale" into this paper per arXiv moderation: the machine-checked action-safety core preserved across six further releases, six new invariant families with teeth, and real-hardware self-improvement; suite grown from 122 to 563 tests. Software, gate suite, run artifacts, and TLA+ spec are open source