AI 中文总结
ChurnBench通过时间线式数据织物和真实值账本,证明刷新调度而非缓存年龄决定智能体AI的陈旧性,TTL配置是漂移基准的关键变量。
AI 中文摘要
在生产环境中,智能体系统需要回答关于分布在多个位置且持续变化的数据的问题:许可证被重新分配、用户被移除、价格发生变化、合同被续签。现有的检索基准将数据冻结,因此它们无法判断智能体的答案是否仍然正确,只能判断是否找到了正确的段落。我们提出了ChurnBench,一个开源基准,它将一个四来源的企业数据织物生成为时间线而非快照。每次更改都写入仅追加的真实值账本,黄金答案根据该账本计算,而非实时存储。因此,当答案在数据检索时正确但在评估时错误,会被检测并标记为新鲜度错误,区别于推理错误;我们通过在报告每个案例的两个时间戳上解析真实值来验证这一点。使用该工具,我们发现当系统按计划刷新时,缓存年龄无法预测陈旧性。在缓存年龄为1、14和28天时,新鲜度错误分别为7、4和4,因为计划刷新通过生存时间限制了陈旧性,且在任何窗口内均未观察到TTL过期。受控消融实验确认了机制:禁用分层刷新在28天时将新鲜度错误从4提高到45,而在1天时保持不变。因此,漂移基准应扫描的变量是针对每个实体变化率的TTL配置,而非漂移窗口长度。ChurnBench、评估工具以及所有逐错误数据均以开源形式发布。
英文摘要
In production, agentic systems answer questions over data that lives in several places and keeps changing: licenses are reassigned, users offboarded, prices changed, contracts renewed. Existing retrieval benchmarks freeze the data, so they cannot ask whether an agent's answer is still true, only whether it found the right passage. We present ChurnBench, an open-source benchmark that generates a four-source enterprise data fabric as a timeline rather than a snapshot. Every change is written to an append-only ground-truth ledger, and gold answers are computed from that ledger, never from the live stores. An answer that was correct when its data was retrieved but wrong when evaluated is therefore detected and labeled a freshness error, distinct from a reasoning error; we validate this by resolving ground truth at both timestamps for every case reported. Using the instrument, we find that when a system refreshes on a schedule, cache age does not predict staleness. Across cache ages of 1, 14, and 28 days, freshness errors were 7, 4, and 4, because scheduled refresh bounds staleness by time-to-live, and no TTL lapse was observed in any window. A controlled ablation confirms the mechanism: disabling tiered refresh raises freshness errors from 4 to 45 at 28 days and leaves them identical at one day. The variable a drift benchmark should sweep is therefore TTL configuration against each entity's rate of change, not drift-window length. ChurnBench, the evaluation harness, and all per-error data are released open source.
Comments8 pages, 2 figures, 5 tables. Code, data, and evaluation harness in github: https://github.com/vsingh45/churnbench