AI 中文总结
本研究提出开源工具factwash,用于捕捉AI将传闻洗成事实的改写,明确否定线索检测F1达0.91,LLM见证者可提升线索检测召回率,在mem0 2.0.7上标记8个模糊传闻写入中的5个。
AI 中文摘要
AI系统不断改写信息:对话变成存储的记忆,文档变成答案。改写可能保留主张,但洗去使其可核查的依据、谁说的、他们的确定程度、何时成立——我们称这种失败为事实清洗(factwashing),并发布factwash,这是一个开源的写入时网关,可确定性地捕捉该问题,带有命名标志和证据,而非大型语言模型(LLM)评判。构建它回答了一个实际问题:何时廉价检查就足够,何时需要模型?决定因素是属性是否有有限的表面线索集合。明确否定线索接近可枚举,因此单词列表即可完成并迁移,在未微调文本上达到0.91的F1值。模糊表述(hedging)和归属有开放式的实现,因此词汇表的召回率稳定在近一半,而一个仅需一个问题的LLM见证者在相同精度下,将线索检测召回率分别提高了17和15个百分点。部署时,该见证者可能只会降低判定,因此它提升的是精度而非覆盖范围。我们在105596个独立标注的句子上测量线索检测。然后,一个盲标注的记忆写入语料库定位了该失败:55%的不良写入来自对话传闻,7%来自商务电子邮件(p < 0.001),因此首次部署的问题不是使用哪个检测器,而是该失败是否发生。在未修改的mem0 2.0.7上,该网关标记了8个模糊传闻写入中的5个。
英文摘要
AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim while washing away what made it checkable, who said it, how sure they were, when it held. We call that failure factwashing, and release factwash, an open-source write-time gate that catches it deterministically, with named flags and evidence rather than an LLM judge. Building it answers a practical question: when does a cheap check suffice, and when do you need a model? What decides is whether the property has a bounded surface-cue inventory. Explicit negation cues are close to enumerable, so a word list finishes and transfers, reaching 0.91 F1 on untuned text. Hedging and attribution have open-ended realizations, so vocabulary plateaus near half recall, and a one-question LLM witness recovers +17 and +15 points of cue-detection recall at equal precision. Deployed, that witness may only lower a verdict, so it buys precision rather than coverage. We measure cue detection on 105,596 independently annotated sentences. A blind-labelled corpus of memory writes then locates the failure: 55% of bad writes in conversational hearsay, 7% in business email (p < 0.001), so the first deployment question is not which detector to use but whether the failure occurs at all. On unmodified mem0 2.0.7, the gate flags 5 of 8 hedged-hearsay writes.
Comments15 pages, 3 figures. Code and data: https://github.com/collapseindex/factwash