程序性记忆在变化中的表现:受控Web任务中的复用与干扰
Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks
浏览论文内容
中文总结 AI 辅助
本研究通过受控Web任务实验,发现程序性记忆在适用性不匹配时可能不产生行为干扰,但需进一步探究触发干扰的条件。
中文摘要 AI 辅助
程序性记忆使语言智能体能够复用成功的例程,但复用预设了存储的例程仍然适用。我们研究了当这一预设被故意违反时会发生什么。本研究结合了来自BrowserGym TimeWarp的回顾性、人工辅助的界面适配案例,以及在合成购物决策上进行的受控冻结记忆比较。在有记录的WebShop V1-V6开发路径中,界面特定代码被适配,而单独存储的高层程序未报告发生变化;该阶段不构成自主记忆智能体评估。在受控阶段,一个早期试点产生了一个任务,在该任务中,两种记忆条件选择了更昂贵的物品,而无记忆条件选择了参考最小值。后续探测未建立重复的行顺序或身份绑定模式。然后,我们测试了四种不匹配形式:数量变化、不同的证据表示、局部与全局优化之间的冲突,以及分布式促销证据,共跨越32个正式单元。每个单元使用一次温度为零的生成,采用相同的本地qwen3:8b配置,且无自适应重试。在这些配对中,当当前任务证据明确且充分时,预定义的诊断性干扰特征均未出现在为其定义的任务上。该结果确定了一个经过测试的非干扰区域:程序性记忆可以不匹配而不产生行为干扰。这并未确立普遍安全性或机制。剩余的问题是,哪些额外条件会将适用性不匹配转化为可观察的、由记忆引起的错误。
英文摘要
Procedural memory lets language agents reuse successful routines, but reuse presumes that a stored routine remains applicable. We study what happens when that presumption is deliberately violated. The study combines a retrospective, human-assisted interface-adaptation case from BrowserGym TimeWarp with controlled frozen-memory comparisons on synthetic shopping decisions. During the documented WebShop V1-V6 development path, interface-specific code was adapted while the separately stored high-level procedure was not reported to change; this phase does not constitute an autonomous memory-agent evaluation. In the controlled phase, an early pilot produced one task on which two memory conditions selected a more expensive item while the no-memory condition selected the reference minimum. Follow-up probes did not establish a recurring row-order or identity-binding pattern. We then tested four forms of mismatch: changed quantities, a different evidence representation, a conflict between local and global optimization, and distributed promotion evidence, across 32 formal cells. Each cell used one temperature-0 generation with the same local qwen3:8b configuration and no adaptive retry. Across these pairs, none of the predefined diagnostic interference signatures appeared on the tasks for which they were defined when current-task evidence was explicit and sufficient. The result identifies a tested region of non-interference: a procedural memory can be mismatched without becoming behaviorally disruptive. It does not establish general safety or a mechanism. The remaining question is which additional conditions turn applicability mismatch into observable, memory-caused error.