arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05906cs.CL

用于反馈驱动智能体修复的因果情景记忆

Causal Episodic Memory for Feedback-Driven Agent Repair

  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
  • Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT), VNU-HCM(胡志明市理工大学计算机科学与工程学院(HCMUT,VNU-HCM))

机构由 AI 辅助整理,请以论文原文为准。

Khang Nhat Hoang Vo, Tam Minh Chu, Anh Trac Duc Dinh, Thuyen Vinh Ha Bui, Tho Quan

AI总结:

本文提出无需训练的MERIT智能体,通过维护双极性记忆辅助LLM智能体修复,在Spider、BIRD数据集上提升了Text-to-SQL任务的执行准确率,同时明确了因果跨查询记忆的适用场景。

AI中文摘要:

修复失败的大语言模型(LLM)智能体常丢弃成功的修正结果,迫使后续情景重新发现类似解决方案。本文研究最终确定的修复结果能否在不更新参数的情况下改善后续的文本到SQL(Text-to-SQL)情景。我们提出MERIT,这是一种无需训练的智能体,它维护一个在线双极性记忆,存储经预言机(oracle)验证的修正结果与观察到的失败方向。在预言机辅助的基准反馈下,只有早期已完成情景的记忆才有资格被检索。一个确定性分类器会分配粗略的失败类型,该类型会在冻结模型生成每次修正前,对混合词汇-密集检索器进行条件设置。使用初始预测和修复预算相同的Qwen2.5-7B-Instruct模型,MERIT在Spider数据集上将执行准确率从66.34%提升至69.79%,在BIRD数据集上从47.35%提升至48.44%。配对分析为Spider的提升提供了明确证据,但在BIRD上的证据较弱。在两个基准上,MERIT与无类型动态检索的性能均未达到可靠区分,而Reflexion式记忆在BIRD上达到51.24%的准确率,但推理成本显著更高。消融实验表明,负向记忆的贡献较小,类型条件设置与词汇-密集排序的价值具有数据集依赖性,而模式局部经验提供最一致的收益。这些结果阐明了因果跨查询记忆何时能改善修复,以及何时更广泛的记忆表征仍更可取。

英文摘要:

LLM agents that repair failures often discard successful corrections, forcing later episodes to rediscover similar solutions. We study whether finalized repair outcomes can improve subsequent Text-to-SQL episodes without parameter updates. We introduce MERIT, a training-free agent that maintains an online dual-polarity memory of oracle-verified corrections and observed unsuccessful directions. Under oracle-assisted benchmark feedback, only memories from earlier finalized episodes are eligible for retrieval. A deterministic classifier assigns a coarse failure type, which conditions a hybrid lexical-dense retriever before the frozen model generates each revision. Using Qwen2.5-7B-Instruct with identical initial predictions and repair budgets, \method{} improves execution accuracy over stateless iterative repair from \(66.34\%\) to \(69.79\%\) on Spider and from \(47.35\%\) to \(48.44\%\) on BIRD. Paired analyses provide clear evidence for the Spider gain but weaker evidence on BIRD. MERIT is not reliably separated from untyped dynamic retrieval on either benchmark, while Reflexion-style memory reaches \(51.24\%\) on BIRD at substantially higher inference cost. Ablations show that negative memory contributes modestly, the value of type conditioning and lexical-dense ranking is dataset dependent, and schema-local experience provides the most consistent benefit. These results clarify when causal cross-query memory improves repair and when broader memory representations remain preferable. Our implementation is available here:

↑