arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

D²ACCI:一种用于保留证据的智能体记忆的双循环诊断协议

D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory

Xule Liu, Yijun Liu, Chao Li, Shao Kun

arXiv 2608.17756首次发表:更新:

AI 中文总结

该研究针对LLM智能体持久记忆流程故障难以定位的问题,提出双循环诊断协议D²ACCI,引入DCR指标与D²ACCI-Eval人工制品,在三个公开基准上取得明确性能增益,填补了记忆系统迭代缺乏可追踪证据的空白。

AI 中文摘要

记忆是大型语言模型(LLM)智能体的关键能力,持久记忆可跨会话扩展该能力,实现回忆、修订与个性化。但其多阶段流程(摄入、检索、过滤、生成)导致故障难以定位:端到端评估仅能发现错误,无法确定是哪个阶段引发的。现有评估常报告整体性能,却缺乏配对统计比较、切片级非回归检查或阶段级诊断追踪。我们提出D²ACCI(诊断驱动的基于人工制品的闭环受控迭代),这一双循环协议的外层诊断门控基于配对证据、受保护切片监控与追踪级可定位性,对记忆干预措施进行推进、特征标记或拒绝。我们还引入DCR,一种衡量故障是否仍可定位的分级可观测性指标,以及D²ACCI-Eval,一种用于门控重放的可复用人工制品。我们在MemStack中实例化该协议,并在三个公开基准上评估,在LoCoMo上达到93.59%,在LongMemEval上达到90.93%,在PersonaMem-V2上达到57.20%。五次配对 ablation 显示,补充提取、会话记忆检索和Forget Guard可产生统计显著的增益(+1.9至+3.7个百分点,所有p值≤0.003)。相比之下,BM25/RRF被保留为受监控的特征标记——这一区别是仅靠整体评估无法察觉的。诊断审计显示,丰富的追踪比仅结果重标记大幅提升了根本原因一致性。诊断人工制品的DCR@3达到98-100%,而仅结果日志的DCR@3为0%。这些结果表明,稳健的记忆系统迭代需要可追踪、基于统计且感知回归的证据——这正是D²ACCI所填补的空白。

英文摘要

Memory is a key capability of LLM agents. Persistent memory extends this across sessions---enabling recall, revision, and personalization. Yet its multi-stage pipeline (ingestion, retrieval, filtering, generation) makes failures difficult to localize: end-to-end evaluation reveals that an error occurred, but not which stage caused it. Existing evaluations often report aggregate performance without paired statistical comparisons, slice-level non-regression checks, or stage-level diagnostic traces. We propose D$^2$ACCI (Diagnostic-Driven Artifact-based Closed-loop Controlled Iteration), a dual-loop protocol whose outer diagnostic gate promotes, feature-flags, or rejects memory interventions based on paired evidence, protected-slice monitoring, and trace-level localizability. We further introduce DCR, a graded observability metric that measures whether failures remain localizable, and D$^2$ACCI-Eval, a reusable artifact for gate replay. We instantiate the protocol in MemStack and evaluate on three public benchmarks, achieving 93.59% on LoCoMo, 90.93% on LongMemEval, and 57.20% on PersonaMem-V2. Five paired ablations show that supplement extraction, session-memory retrieval, and Forget Guard yield statistically significant gains (+1.9 to +3.7pp, all p $\le$ .003). In contrast, BM25/RRF is retained as a monitored feature flag---a distinction invisible to aggregate-only evaluation. A diagnostic audit shows enriched traces substantially improve root-cause agreement over result-only relabeling. Diagnostic artifacts reach 98--100% DCR@3 versus 0% for results-only logs. These results establish that robust memory-system iteration demands traceable, statistically grounded, and regression-aware evidence---exactly the gap D$^2$ACCI fills.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑