arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

继承式智能体内存预算验证中的计划指针与记录指令形式

Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory

Kazuki Nakayashiki

arXiv 2609.03450首次发表:更新:

发表机构

Glasp(Glasp)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究在智能体内存预算验证中,通过12项注册实验分析计划指针、记录准则等编辑对模型选择的影响,明确了不同模型及设置下的效果差异。

AI 中文摘要

继承了6条单行记忆的智能体在行动前最多可调取1条存档源记录;存储中写入的指令可引导该选择:指向记录的指针、识别该记录的准则,或两者结合。在对1个工具谱系开展的12项注册研究(共14760次尝试)中,我们测量了每种形式下请求的去向。在6个直接提供的模型上,长度匹配的准则比单纯的ID高出35.0个点[+31.2, +38.8](研究D);在9个模型组成的OpenRouter服务面板上,该对比未通过注册的优越性规则(研究E)。在3个Claude模型上,追加ID会抵消准则的作用(Opus 5:40/40变为0/40;研究F-x);6个字节匹配的编辑使每个精确字符串产生各自的效果(研究G),且每个单元80次重复的重运行使30项重复对比中的15项在误差范围内,15项未解决,无项超出(研究G')。批准线(Opus 5上+96.0个点)和2个信用的预算在3个模型上均恢复了目标(研究J);在5个准则字符串中,后缀的抵消作用对Opus 5的5种措辞中的4种成立,对Fable 5.1的5种措辞全部成立(研究H2);在第二个存储中,所有模型均遵循准则(研究H1)。若延续至决策,准则使选择偏向当前记录(Opus 5上+100.0个点),而在Fable 5.1上则偏离当前记录(研究I)。1个字符的计划指针的效果(+78.0个点;研究B,在首次仓库报告修正后)在预注册的重运行中返回相同结果(+81.7个点;研究B')。所有结果均为固定面板上精确编辑的描述性效果,带有注册区间,且不涉及机制主张。

英文摘要

A model that inherits one-line memories may pull one archived source record before acting; a directive in the store can steer that pull: a pointer, a criterion or both. Across sixteen registered studies (179,352 attempts) we measured where the request goes under each form; every result is descriptive, with registered intervals, no mechanism claim. A length-matched criterion exceeded a bare id on six direct-provider models (D) and failed its registered superiority rule on a nine-model OpenRouter panel (E). On generated worlds (K2-K5): the two registered signatures held on Opus 5 and Fable 5.1, Fable 5 followed the same sign, Haiku 4.5 reversed, and Sonnet 5, the GPT-5.6 endpoints and GPT-6 Astra lay near zero (K2). With a defensive adapter at five gains, the 70B rule for a gain-dependent change of the composite - criterion contrast was not met (K3 and K4); under the 8B attenuation rule (0.95 intervals: slope below zero; change beyond the margin), the 8B change of -17.5 [-26.7, -8.1] did not meet it on 36 families (K4) and at registered power on 337 families -16.6 [-19.4, -13.8] did (realised one-sided error at the margin 1.8 to 3.2% per corner of a finite grid, nominal 2.5%, not a uniform-error guarantee; K4's status stands; K5, first ladder), while a second SecAlign++ adapter under the imposed Meta-SecAlign template did not (-11.8 [-14.3, -9.3]; K5, second ladder); no NOT-MET is a statement that the contrast was unchanged; their difference (+4.7 [+2.3, +7.2]) describes two fixed execution paths, licenses no superiority, equivalence or 'significant difference' claim; nothing follows from the statuses differing (K5). Intervals describe family-reweighting stability conditional on the execution, not reproducibility across engine executions; audit replays were neither substituted for nor averaged into outcomes; no missingness gate fired and directional completions changed no status.

Comments65 pages, 7 figures, 44 tables. Sixteen registered studies (179,352 attempted episodes) on one instrument lineage; every package was frozen, hashed and externally deposited before its first confirmatory call. v2 adds Studies K2-K5 (generated worlds; a defensive adapter at five gains). Manuscript, source, records and the generator of every number: doi:10.5281/zenodo.22267220

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑