MemoryAthena:潜在记忆与生成记忆上的自适应路由
MemoryAthena: Adaptive Routing over Latent and Generated Memories
- ELLIS Institute of Finland(芬兰ELLIS研究所)
- University of Turku(图尔库大学)
- University of Technology Sydney(悉尼科技大学)
- University of Science and Technology of China(中国科学技术大学)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MemoryAthena 通过轻量级因果路由头学习何时用生成记忆(GE/GH)修正直接检索(E),在问答和通用 NLP 任务上分别提升平均分至 39.28 和 79.13。
AI中文摘要:
学习型记忆方法将信息存储在一个显式表格中,并通过一个独立的读取器来消费这些信息,从而允许寻址、存储和读取被独立地修改。我们研究有用的记忆是否也可以被生成而不仅仅是检索。MemoryAthena 使用三条通路:直接的 Engram 检索(E)、从检索到的 Engram 线索生成(GE)、以及在不查阅记忆表格的情况下从因果骨干状态生成(GH)。生成的记忆是有条件地有用的:它可以在一种情境中补充 E,但在另一种情境中干扰 E。因此,MemoryAthena 将 E 视为锚点,并学习何时生成的表示应该介入。在骨干、记忆、生成器和读取器均被冻结的情况下,一个轻量级的因果路由头根据 GE 和 GH 相对于 E 的反事实未来令牌似然优势进行训练。在推理时,一个被接受的候选通过有界插值修改 E 残差,而拒绝则精确地恢复直接通路。在问答任务上,MemoryAthena 将同一检查点的直接通路的五个任务平均分从 37.65 提高到 39.28,而六个任务的一般 NLP 平均分从 76.73 提高到 79.13。完整的记忆侧系统包含约 201M 参数,不包括冻结的骨干。进一步的分析显示 E、GE 和 GH 在不同任务和输入上具有互补的优势。这些结果支持生成的记忆作为对直接检索的选择性修正,并强调路由何时、何种以及多强地介入是核心挑战。
英文摘要:
Learned-memory methods store information in an explicit table and consume it through a separate reader, allowing addressing, storage, and reading to be modified independently. We study whether useful memory can also be generated rather than only retrieved. MemoryAthena uses three pathways: direct Engram retrieval (E), generation from retrieved Engram cues (GE), and generation from causal backbone states without consulting the memory table (GH). Generated memory is conditionally useful: it can complement E in one context but interfere with it in another. MemoryAthena therefore treats E as an anchor and learns when a generated representation should intervene. With the backbone, memory, generators, and readers frozen, a lightweight causal routing head is trained from counterfactual future-token likelihood advantages of GE and GH relative to E. At inference time, an admitted candidate modifies the E residual through bounded interpolation, while rejection recovers the direct pathway exactly. On question answering, MemoryAthena raises the five-task average from 37.65 to 39.28 over the direct pathway of the same checkpoint, while the six-task general-NLP average increases from 76.73 to 79.13. The complete memory-side system contains approximately 201M parameters, excluding the frozen backbone. Further analyses show complementary strengths among E, GE, and GH across tasks and inputs. These results support generated memory as a selective correction to direct retrieval and highlight routing when, which, and how strongly to intervene as the central challenge.