一致性门控:通过自一致性准入控制防止大语言模型智能体中的内存污染
ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control
- Florida State University(佛罗里达州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究大语言模型智能体内存污染问题,提出一致性门控方法,在提交候选事实前向模型查询K次获软支持分数,超阈值才准入。该机制与模型无关、无需微调,经实验在四个模型主干上减少了污染,还发布了相关基准和实现。
AI中文摘要:
在多轮运行的大语言模型智能体中,会在外部内存存储中积累事实并将其用作下游推理的前提。然而,一步中产生的幻觉事实会持续作为后续步骤的错误前提,即内存污染问题。现有内存管理解决了检索和容量问题,但未解决写入时的正确性。我们提出了一致性门控,这是一种写入时的准入门控,在提交从上下文提取的候选事实前,向大语言模型查询K次以获得软支持分数,仅当平均值超过阈值时才准入该事实。该机制与模型无关,无需微调,在延迟敏感部署的对数概率变体中简化为单次前向传递。为衡量对自然数据的影响,我们构建了两个真实对话基准和一个结构化合成语料库。在四个大语言模型主干上,一致性门控相对于写入所有内容的基线减少了每个基准上的污染,成本集中在源上下文中仅隐含陈述的事实上。我们发布了所有三个基准以及门控实现。
英文摘要:
LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A hallucinated fact written at one step therefore persists as a false premise for every subsequent step, a failure mode we call memory contamination. Existing memory management addresses retrieval and capacity but not write-time correctness; this admission problem cannot be solved by utility- or recency-based criteria, and uncontrolled contamination compounds across long trajectories. We propose ConsistencyGate, a write-time admission gate that, before committing a candidate fact m extracted from context c, queries the LLM K times for a soft support score and admits m only when the average exceeds a threshold. The mechanism is model-agnostic, requires no fine-tuning, and reduces to a single forward pass in a log-probability variant for latency-sensitive deployments. To measure the effect on natural data, we construct two real-conversation benchmarks (LoCoMo-Contam and MSC-Contam) by planting controlled single-detail corruptions in long-term conversations from LoCoMo and MSC, and complement them with a structured synthetic corpus (MemContam) that isolates a near-oracle upper bound. Across four LLM backbones, ConsistencyGate reduces contamination on every benchmark relative to a write-everything baseline, with the cost concentrated on facts that are stated only implicitly in the source context. We release all three benchmarks together with the gate implementation.