超越端点分数:持续知识更新的时间与容量条件评估
Beyond Endpoint Scores: Time- and Capacity-Conditioned Evaluation of Continual Knowledge Updating
浏览论文内容
中文总结 AI 辅助
该研究指出持续知识更新方法的排名受评估时间和重放适应容量共同影响,提出需报告轨迹与容量扫描以确定稳健胜者,发现周期性层次结构是更新成本更低的操作点。
中文摘要 AI 辅助
持续知识更新方法通常仅通过一个最终检查点和一个常规适配器秩就被宣称更优,但我们表明这不足以识别更优的操作点。在固定周期性层次结构的情况下,我们在24个月的Wikidata流上,通过变化评估月份、重放LoRA秩和查询表述,将其与累积重放进行比较。在该范围内,明显的胜者会发生变化:在Qwen2.5-1.5B上,层次结构相对于秩8重放的5.0点优势,在秩72重放时变为11.6点劣势;在高秩下,与对齐的端点整合可表明平局,而时间平均重放则领先9-13点。在Llama-3.2-1B和保留的转述上也出现了相同的秩条件反转。这些结果表明,持续更新中的方法排名可同时取决于性能测量时间和基线接收的重放侧适应容量。因此,我们建议报告轨迹和容量扫描,并仅当排序在评估范围内稳定时才声明稳健胜者;否则,比较应报告胜者区域和保留-稳定性-成本前沿。根据该协议,周期性层次结构是更新成本更低的操作点,而非质量胜者。
英文摘要
Continual knowledge-updating methods are often declared superior from one final checkpoint and one conventional adapter rank. We show that this can be insufficient to identify the better operating point. Holding a periodic hierarchy fixed, we compare it with cumulative replay over a 24-month Wikidata stream while varying evaluation month, replay LoRA rank, and query formulation. The apparent winner changes across this region: on Qwen2.5-1.5B, the hierarchy's 5.0-point advantage over rank-8 replay becomes an 11.6-point deficit against rank-72 replay, and at high ranks a consolidation-aligned endpoint can suggest a tie while time-averaged replay leads by 9-13 points. The same rank-conditioned reversal appears on Llama-3.2-1B and held-out paraphrases. These results show that method ranking in continual updating can depend jointly on when performance is measured and how much replay-side adaptation capacity the baseline receives. We therefore propose reporting trajectories and capacity sweeps, and declaring a robust winner only when the ordering is stable across the evaluation region; otherwise, comparisons should report winner regions and retention-stability-cost frontiers. Under this protocol, the periodic hierarchy is a lower-update-cost operating point, not a quality winner.
发表机构
- Yonsei University(延世大学)
- Sooil Development
- NCreate
机构由 AI 辅助整理,请以论文原文为准。