AI 中文总结
该研究指出当前知识编辑基准无法衡量SERAC系列的范围决策,以INLAY为案例验证其结构性缺陷,还披露了INLAY的不足及自身路由机制的两个未影响核心结果的错误。
AI 中文摘要
SERAC系列的每一种基于记忆的知识编辑器都依赖于范围决策:给定一个查询,存储的编辑是否适用?我们报告称,当前的知识编辑基准完全无法衡量这一决策。我们构建了INLAY(一种无梯度编辑器,其模型处于冻结状态,编辑内容存储在外部可寻址内存中,应用编辑是解码时沿某个token的未嵌入方向添加偏置)以获取精确的每查询真实值,随后在涵盖三个数据集、三种输入条件的1689个查询上执行所有候选路由器动作。在所有9个数据集-条件组合中,每次选择最佳动作的神谕路由器与一行静态策略的性能持平,达到四位小数;任何每查询路由方法可获得的最大增益为0.00点;在1689次查询中,弃权(不执行)动作从未成为唯一获胜动作。其原因是结构性的:这些是反事实基准,其评估问题要求给出编辑后的答案,因此根据构造,从参数知识中获取答案是错误的,而没有负样本的基准无法奖励分类器的拒绝能力。这一问题不仅存在于我们的系统,还延伸到基准所评估的整个范围分类器家族。我们直接验证了这一机制:通过对一半样本将查询自身的编辑从索引中 withheld(保留),将合并的性能空间从恰好+0.0000提升至+0.0420,并使弃权(不执行)动作首次获胜。我们还报告了INLAY未取得优势的情况(WISE在Qwen2.5-7B CounterFact上优于它,且在严格匹配的RippleEdits上,检索增强生成优于包括INLAY在内的所有测试方法),并披露了在对自身路由机制进行自我审计时发现的两个错误,这两个错误均未改变已发布的标题数字,其变化在噪声范围内。
英文摘要
Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query, does a stored edit apply? We report that current knowledge-editing benchmarks cannot measure this decision at all. Using INLAY, a gradient-free editor we built to obtain exact per-query ground truth (the model is frozen, edits live in an external addressable memory, and applying an edit is a bias added along one token's unembedding direction at decode time), we execute every candidate router action on 1,689 queries spanning three datasets and three input conditions. An oracle router choosing the best action every time ties a one-line static policy to four decimal places in all nine dataset-by-condition cells: the maximum attainable gain of any per-query routing method is 0.00 points. Abstention is the sole winning action zero times out of 1,689. The cause is structural: these are counterfactual benchmarks whose evaluation question asks for the post-edit answer, so answering from parametric knowledge is wrong by construction, and a benchmark without negatives cannot reward a classifier's ability to reject. This generalizes beyond our system to the whole scope-classifier family the benchmarks are used to evaluate. We confirm the mechanism directly: constructing the missing condition ourselves, by withholding a query's own edit from the index for half the sample, moves pooled headroom from exactly +0.0000 to +0.0420 and gives abstention its first wins. We also report where INLAY itself does not win (WISE beats it on Qwen2.5-7B CounterFact, and retrieval-augmented generation beats every method we tested, INLAY included, on rigorously matched RippleEdits), and disclose two bugs found during a self-audit of our own routing machinery, neither of which changed a published headline number outside noise.
Comments12 pages, 5 figures, 6 tables. Code and data: https://github.com/Aditya-PS-05/INLAY