AI 中文总结
本研究提出配对回放协议,通过时点发现与时间违规率审计,揭示历史语料库回放中记忆泄漏导致评估虚高,并强调历史评估需同时版本化记忆与语料库。
AI 中文摘要
离线回放应估计发现系统在历史时点可能检索到的内容,然而冻结语料库却使交互记忆不受约束。我们通过历史状态 $(D_t, \ heta_t, M_{< i})$ 形式化时点(PIT)发现,并引入一种仅改变记忆可用性的配对回放。该协议从仅行为轨迹构建PIT和全流Future视图,并使用时间违规率审计所选条目。在三个表-文本领域、两种流模式、两个检索器和五个随机种子(216,000行)中,Future将资产召回率@100提升了2.62-5.24个百分点;所有12个配对区间均排除零。在仅行为轨迹记忆下,PIT的表现不如无记忆的Stateless条件;Future掩盖了该损害的32.7-48.4%。对于模拟的正反馈缓存,PIT比Stateless高出4.65-18.96个百分点,而Future又额外高出4.11-9.58个百分点。在五个带时间戳的FreshStack主题上,Future超过PIT 2.72个百分点[1.75, 3.71]。历史评估必须将记忆与语料库一同进行版本化并验证。
英文摘要
Offline replay should estimate what a discovery system could retrieve at a historical point, yet freezing the corpus leaves interaction memory unconstrained. We formalize point-in-time (PIT) discovery through historical state $(D_t, θ_t, M_{< i})$ and introduce a paired replay that changes only memory availability. The protocol constructs PIT and full-stream Future views from behavior-only traces and audits selected entries with a Temporal Violation Rate. Across three table-text domains, two stream regimes, two retrievers, and five seeds (216,000 rows), Future inflated Asset Recall@100 by 2.62-5.24 points; all 12 paired intervals excluded zero. With behavior-only trace memory, PIT underperformed the no-memory Stateless condition; Future masked 32.7-48.4% of that harm. For a simulated positive-feedback cache, PIT added 4.65-18.96 points over Stateless while Future added another 4.11-9.58 points. On five timestamped FreshStack topics, Future exceeded PIT by 2.72 points [1.75, 3.71]. Historical evaluation must version and validate memory with the corpus.
Comments12 pages, including figures and tables