发表机构
International Digital Economy Academy (IDEA)(国际数字经济研究院(IDEA))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过大规模迁移实验发现,修复失败回合能提升相关任务性能,但修正记忆的价值独立于复用经验,需同时参考旧版本与全新起点。
AI 中文摘要
修复一个回合是否会让其经验成为下一个任务更好的记忆?我们将同一失败源在修复被接受前后迁移到固定目标,并同时进行独立执行。我们的3300次运行覆盖100个ThinkingBox配对和相同的100个APEX配对,在有和没有源状态继承的情况下,共在11种条件下进行。ThinkingBox的Full/Skill/Hybrid修正增益分别为44/29/32个百分点,修正后的性能比独立执行高25/22/18个百分点;在任务族层面推断会减弱。然而,Full相对于Skill的15个百分点更大的修正差距中,有12个百分点来自更差的未修正性能,而非更好的修正记忆。此外,Full的46次向上转换中有22次恢复了观察到的基线成功。APEX的两种机制均未建立可比的总体修正收益。行动证据将工作流增益与可复用义务联系起来,将约定冲突与源局部选择联系起来。文本APEX的接受执行达到52%,而其摘要为40%,没有稳健的全局/组级优越性或相对于独立执行的既定优势。较小的交接减少输入但增加调用。因此,修复经验的价值不同于复用经验的价值:记忆更新需要先前的版本参考和全新的参考。
英文摘要
Does repairing an episode make its experience a better memory for the next task? We transfer the same failed source before and after accepted repair to a fixed target, alongside independent execution. Our 3,300 runs cover 100 ThinkingBox pairs and the same 100 APEX pairs with and without source-state inheritance, under eleven conditions. ThinkingBox's Full/Skill/Hybrid correction gains are 44/29/32 percentage points, with corrected performance 25/22/18 points above independence; inference weakens at the task-family level. Yet 12 of Full's 15-point larger correction gap over Skill come from worse uncorrected performance, not better corrected memory. Moreover, 22 of Full's 46 upward transitions restore observed baseline success. Neither APEX regime establishes comparable aggregate correction benefits. Action evidence connects workflow gains with reusable obligations and convention conflicts with source-local choices. Text APEX's accepted execution reaches 52% versus its summary's 40%, without robust global/group-level superiority or an estab- lished advantage over independence. Smaller handoffs reduce input but increase calls. The value of repairing experience is therefore distinct from the value of reusing it: memory updates require both a previous-version reference and a fresh-start reference.