arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当记忆更新但行为未更新时:修复个性化智能体响应中的隐式陈旧依赖

When Memory Updates but Behavior Does Not: Repairing Implicit Stale Dependencies in Personalized Agent Responses

Haofei Sun, Lin He

arXiv 2608.01619首次发表:更新:

AI 中文总结

该研究针对智能体存在的隐式陈旧依赖问题,提出StateAuditor反向审计方法,在STALE基准上实现显著性能提升,验证了转换机制对修复IPA缺口的有效性。

AI 中文摘要

具备记忆增强功能的智能体可能知晓用户存储的状态已过时,却仍基于旧值进行规划,STALE基准将此称为隐式策略适配(IPA)缺口。我们确定了一个结构性成因:基于草稿的验证机制仅检查响应内容,而在开放式响应中,陈旧依赖通常未被明确表述。因此StateAuditor采用反向审计,从存储状态指向草稿。大型语言模型(LLM)会基于带时间戳的证据提出候选新旧转换;确定性代码将每个引用绑定到单一条目,验证新证据确实更新,且仅允许这些经验证的转换触发修复。验证的是来源和时间顺序,而非语义替代。在STALE的完整协议(400个场景、50个会话历史、每个查询对应一个独立响应)下,严格单查询VTA得分达.736,而在相同评估者下,我们的锁定前体模型得分为.686,配对增益为+5.0个百分点(95%置信区间[+2.9, +7.2]),该增益几乎全部来自IPA和前提抗性(PR)。来自第三方模型家族的基准自身评估者也重现了这一增益(.738对比.680)。在独立的跨家族偏好演化基准HorizonBench上,基于黄金衍生结构化存储的完整草稿-审计-修复流程提升了当前偏好准确率(用户聚类p<.01),不过匹配对照显示,大部分外部增益来自草稿侧审计本身;更困难的创作生命周期集未产生增益,在控制错误无效的同时限定了主张范围。相反,在STALE上,匹配对照(相同证据、适配器和调用预算)仅得.692(较前体+0.6,无统计学显著性),这将STALE的增益归因于转换机制,而非新增上下文或调用。我们不主张通用智能体记忆的相关结论。

英文摘要

Memory-augmented agents can know that a user's stored state is outdated and still plan around the old value. The STALE benchmark calls this the implicit policy adaptation (IPA) gap. We identify one structural contributor: draft-anchored verification checks what a response says, and in an open-ended response the stale dependency is usually unsaid. StateAuditor therefore audits in the opposite direction, from stored state to draft. An LLM proposes candidate old-to-new transitions from timestamped evidence; deterministic code pins each quotation to a single entry, checks that the new evidence really is newer, and lets only these verified transitions trigger repair. What is verified is provenance and chronology - not semantic supersession. On STALE's full protocol (400 scenarios, 50-session histories, one independent response per query), strict single-query VTA scores .736 against .686 for our locked predecessor under the same judge: a +5.0-point paired gain (95% CI [+2.9, +7.2]) coming almost entirely from IPA and premise resistance (PR). The benchmark's own judge, from a third model family, reproduces the gain (.738 vs. .680). On an independent cross-family preference-evolution benchmark (HorizonBench), the full draft-audit-repair pipeline over a gold-derived structured store raises current-preference accuracy (user-clustered p<.01), though a matched control shows most of this external gain is the draft-side audit itself; a harder authored lifecycle set gives no gain, bounding the claim while false invalidation stays controlled. On STALE, by contrast, a matched control (same evidence, adapter, and call budget) scores only .692 (+0.6 over the predecessor, n.s.), attributing the STALE gain to the transition machinery rather than added context or calls. We make no claim about general-purpose agent memory.

Comments9 pages, 4 figures; supplementary material in ancillary files

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑