自清洁仍被捕获:一个用于衡量智能体自写存储中错误的原始度量,以及分数下降实际衡量的内容
Self-Cleaning and Captured Anyway: One Measured Primitive for Error in a Store an Agent Writes to Itself, and What a Falling Score Actually Measures
浏览论文内容
中文总结 AI 辅助
本研究提出一个无参数的复制函数原始度量,证明智能体自写存储的误差存在硬上限,该度量能高精度预测漂移方向,且规模无法避免捕获。
中文摘要 AI 辅助
一个智能体将其结论写入一个它之后检索的存储中,闭合了一个通常被报告为单向污染的循环。将循环推向无限任期极限并针对仅追加存储,会呈现不同的图景:由于写入从不删除,可达状态空间存在一个硬上限 (n-1)/n,因此结果是两个边缘之间的选择而非衰减。在 f_0 = 0.9 时,两种模式之间的区间在 220 次运行中占 3.6%,而均匀分布本应占 20.6%,且在前 15 次运行中严格为空;合并均值描述了其所汇总运行的 8.2%,中位数描述了 68.2%。模型贡献的一切都由一个无拟合参数的原始度量承载,即复制函数 γ(φ):在 36 个 Wikidata 事实上,sign(γ̂ - γ_crit)(其中 γ_crit = 1/k,r = 0,w = 1)在 360 次真实事实运行中预测了 353 次的漂移方向(同一批次中 40 次合成运行中的 39 次)。规模无法拯救存储:合并前沿捕获率为 0.850,其中 claude-sonnet-4.5 在 20 个种子中的 20 个上被捕获,而我们注册的预测为 <0.5。区间测试的是可区分性而非计数:在真实事实上,多值运行的占用率是其余部分的 6.4 倍。它在 f_0 ∈ {0.1, 0.3, 0.5} 时存活,捕获在 f_0 = 0.5 时达到峰值,并且在四项干预中(标准首先冻结),在匹配预算下时序主导比例,而一致性门将每个模型驱动至 0.993。重采样单元是种子,在合并水平上的设计效应为 3.75:在 44 个种子的对照下,支持声明 4 的排序从三个种子时的 Spearman +0.98 崩溃至四十四个种子时的 +0.31-0.80,而声明 2 的排序在那里是精确的(+1.00,p = 0.017)。所有 87 个分级行都在附录 W 中,其中 37 个被分级为撤回、失败、自我纠正、不可判定或已承认的限制,而 50 个不是。
英文摘要
"An agent that writes its conclusions into a store it later retrieves from closes a loop usually reported as one-way contamination. Taking the loop to the infinite-tenure limit against an append-only store gives a different picture: because writing never deletes, the reachable state space has a hard upper edge at (n-1)/n, so the outcome is a choice between two edges rather than a decay. At f_0 = 0.9 the interval between the two modes holds 3.6% of 220 runs where a uniform spread would put 20.6%, and is strictly empty on the first 15; the pooled mean describes 8.2% of the runs it summarises, the median 68.2%. Everything the model contributes is carried by one measured primitive with no fitted parameter, the copy function γ(ϕ): on 36 Wikidata facts, sign(\hatγ - γ_{crit}), with γ_{crit} = 1/k at r = 0, w = 1, predicts the direction of drift on 353 of 360 real-fact runs (39 of 40 synthetic in the same batch). Scale does not rescue the store: pooled frontier capture is 0.850, with claude-sonnet-4.5 captured on 20 of 20 seeds against our registered prediction of <0.5. What the interval tests is distinguishability rather than count: on the real facts, multi-valued runs have 6.4x its occupancy of the rest. It survives at f_0 in {0.1, 0.3, 0.5}, capture peaks at f_0 = 0.5, and of four interventions with criteria frozen first, timing dominates fraction at matched budget while a consistency gate drives every model to 0.993. The resampling unit is the seed, at a design effect of 3.75 on a pooled level: under a 44-seed control the ordering supporting claim 4 collapses from Spearman +0.98 at three seeds to +0.31-0.80 at forty-four, while claim 2's ordering is exact there (+1.00, p = 0.017). All 87 graded rows are in Appendix W, 37 of them graded withdrawn, failed, self-correcting, undecidable or an acknowledged limit, against 50 that are not."
发表机构
- University of Macau(澳门大学)
- South China University of Technology(华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。