发表机构
Meta AI(Meta AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对智能体记忆保留,提出按语义类别条件化置信度阈值的方法,解决全局阈值无法区分弱支持断言的缺陷,实验显示可减少无支持保留并提升覆盖率。
AI 中文摘要
持久化的智能体记忆,其可靠性取决于保留决策:一个仅由来源弱支持的断言,可能被存储并在之后被当作既定事实重用。我们研究保留决策是否应由一个基于断言语义类别条件化的置信度阈值来指导,而非由单一全局阈值决定,即对证据充分的类别宽松保留,而在推理不可靠的类别上更激进地弃权(不执行)。我们在一个部署的冷启动记忆流水线上,对100个合成人物进行了评估。实证评估源于一个显著的可信度不对称:在4,715个候选断言中,仅有77.9%的价值观和信念类断言得到其来源支持,而所有其他类别的这一比例为96.2%。全局置信度阈值无法区分这两类:它要么接纳无支持的价值观声明,要么丢弃证据充分的声明。将阈值按类别条件化解决了这一权衡。在重复的留出评估中,仅对价值观类设置更严格的阈值,就将无支持的保留从6.2%降至4.0%(相对减少约36%,虽适度但在各折中一致),并且作为佐证,在可比保留率下,相比全局阈值,保留了估计多13个百分点的覆盖率(95%置信区间9.8至16.0)。我们的结果表明,可靠的保留取决于断言的类型,而非仅凭置信度,且类别条件化阈值可以在写入边界充当一种简单有效的选择性预测形式。
英文摘要
Persistent agent memory is only as reliable as its retention decision: an assertion weakly supported by its source can be stored and later reused as established fact. We study whether the retention decision should be governed by a confidence bar conditioned on the semantic category of the assertion rather than by a single global threshold, retaining well-evidenced categories liberally while abstaining more aggressively where inference is unreliable. We evaluate this in a deployed cold-start memory pipeline on 100 synthetic personas. The empirical evaluation is motivated by a sharp reliability asymmetry: across 4{,}715 candidate assertions, only 77.9\% of value and belief assertions are supported by their source, versus 96.2\% for all other categories. A global confidence threshold cannot separate these: it either admits unsupported value claims or discards well-evidenced ones. Conditioning the threshold on category resolves the tradeoff. In repeated held-out evaluation, a stricter bar on values alone reduces unsupported retentions from 6.2\% to 4.0\% (an ${\approx}36\%$ relative reduction, modest but consistent across folds) and, as corroborating evidence, preserves an estimated 13 percentage points more coverage (95\% CI 9.8--16.0) than a global threshold at comparable retention. Our results suggest that reliable retention depends on the type of assertion, not on confidence alone, and that a category-conditioned threshold can act as a simple, effective form of selective prediction at the write boundary.
Comments4 pages, 1 Figure, Accepted to NeurIPS 2026 Social Agent Workshop (https://openreview.net/group?id=NeurIPS.cc%2F2026%2FWorkshop%2FSocialAgent#tab-your-consoles)