arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12476cs.AI

受控持久内存:长程智能体的源绑定状态语义与故障关闭式释放

Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents

Guodong Xu

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对长程智能体内存的矛盾记录问题,提出受控持久内存(GPM)模型,经多组基准测试与服务评估,其在GPM-ReleaseBench、中英文指令分支等场景均实现零不匹配,修复大量基线故障且无退化。

中文摘要 AI 辅助

长期智能体内存通常被视为选择-存储-检索流程,但检索环节无法判定矛盾、被取代、被撤回、被删除或过时的记录是否能支持输出主张。我们提出受控持久内存(Governed Persistent Memory,GPM),这是一种可审计的双时态状态转换模型,具备源绑定准入、派生生命周期状态、当前公共屏障以及故障关闭式结构化释放特性。5个可执行子句涵盖账本完整性、源绑定、冲突隔离、撤回或删除后不可恢复,以及在一个经验证的头部上对新鲜视图的精确主张关闭。在预设的哈希冻结的3600例GPM-ReleaseBench测试中,GPM匹配所有完整结果;3个刻意设计的简单完整策略中最强者匹配1800/3600结果,且在50%的违规案例中出现不匹配的释放。一项独立的密封端到端服务评估在8个查询族中执行实际的录入与释放操作。在其公开披露的V3分支中,受控路径在2400/2400集群上正确,而非受控的本地Qwen2.5-7B为600/2400;它修复了全部1800个基线故障且无退化(单侧95%下界分别为99.875%和99.834%)。后续的V5密封版本覆盖中英文指令分支,带有生成日期固定功能且无冻结后缩减器修正,再次在每个分支上获得2400/2400的正确率。一个独立于生产代码的有限模型探索了331776个语义状态和1990656个查询状态,未发现完整契约反例;10万条轨迹的三引擎差分测试产生零不匹配。这些是有界契约与实现结果,并非开放世界模型准确率或世界真实性的证据。密封服务评估中的受控答案是确定性服务输出;7B结果为非受控对比,并非声称语言模型本身变得完全准确。

英文摘要

Long-term agent memory is usually treated as select--store--retrieve, but retrieval does not decide whether contradictory, superseded, retracted, deleted, or stale records may support an outgoing claim. We introduce Governed Persistent Memory (GPM), an auditable bitemporal state-transition model with source-bound admission, derived lifecycle state, current public barriers, and fail-closed structured release. Five executable clauses cover ledger integrity, source binding, conflict isolation, non-revival after retraction or deletion, and exact claim closure over a fresh view at one verified head. On a prespecified hash-frozen 3,600-case GPM-ReleaseBench, GPM matches all complete outcomes; the strongest of three intentionally simple complete policies matches 1,800/3,600 and makes unmatched releases on 50% of violation cases. A separate sealed end-to-end service evaluation exercises real ingestion and release across eight query families. In its publicly disclosed V3 arm, the governed lane is correct on 2,400/2,400 clusters versus 600/2,400 for ungoverned local Qwen2.5-7B; it repairs all 1,800 baseline failures with no regression (one-sided 95% lower bounds 99.875% and 99.834%). A later V5 reseal over Chinese- and English-command arms, with generation-date pinning and no post-freeze reducer amendment, again obtains 2,400/2,400 per arm. A production-code-independent finite model explores 331,776 semantic and 1,990,656 query states without a full-contract counterexample, and a 100,000-trace three-engine differential yields zero mismatches. These are bounded contract and implementation results, not open-world model accuracy or evidence of world truth. Governed answers in the sealed service evaluation are deterministic service outputs; the 7B result is the ungoverned comparison, not a claim that a language model itself became perfectly accurate.

发表机构

  • Qingdao Guodongxiansheng Network Technology Co., Ltd.(青岛果动先生网络科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑