AI 中文总结
ALLOT提出混合内存路由框架,分离写入优先级与参数预算,在低预算下实现高准确率并减少参数写入,支持按增量价值分配适配能力。
AI 中文摘要
对于大型语言模型(LLMs),当检索已经足够时,参数化适配成本高昂。我们提出ALLOT,一个混合内存路由框架,将学习到的写入优先级与硬性参数预算分离。一个内存感知路由器结合冻结的文本表示、检索置信度和关系元数据;单一排序支持多种写入预算,同时将所有事实保留在外部内存中。在CounterFact数据集上使用Qwen3-4B,ALLOT在20%的参数写入预算下达到0.760的准确率,并恢复了预算匹配的oracle增益的78.4%,相比对每个事实进行双重写入,参数写入减少了80%。在此预算下,将检索和关系特征联合添加到文本中,使归一化oracle增益提高了6.2个百分点。互补的Qwen3-0.6B共享存储结果实现了双重写入级别的准确率,参数写入仅为6-14.5%,跨基准迁移保留了约88%的域内增益。这些结果支持根据增量价值分配适配能力,而不是将每个事实更新视为同等价值的训练目标。
英文摘要
For large language models (LLMs), parametric adaptation is costly when retrieval already suffices. We introduce ALLOT, a hybrid-memory routing framework that separates learned write priority from a hard parametric budget. A memory-aware router combines frozen text representations, retrieval confidence, and relation metadata; a single ranking supports multiple write budgets while preserving all facts in external memory. On CounterFact with Qwen3-4B, ALLOT reaches 0.760 accuracy at a 20% parametric-write budget and recovers 78.4% of the budget-matched oracle gain, with 80% fewer parametric writes than dual-writing every fact. At this budget, jointly adding retrieval and relation features to text improves normalized oracle gain by 6.2 percentage points. Complementary Qwen3-0.6B shared-store results achieve dual-write-level accuracy with 6-14.5% parametric writes, and cross-benchmark transfer retains approximately 88% of in-domain gain. These results support allocating adaptation capacity according to its incremental value rather than treating every factual update as an equally valuable training target.