AI 中文总结
针对外部记忆更新导致多跳问答证据选择困难的问题,提出契约记忆编译器(CMC),先解析更新再遍历关系,在FactConsolidation上达到78.25%准确率,并引入MQuAKE-MemStream数据集。
AI 中文摘要
外部记忆使语言模型智能体能够回答关于历史的问题,这些历史对于答案模型的上下文窗口而言过长。更新带来的问题比检索近期事实更难:改变一个关系可以将多跳问题重定向到关于问题中未提及实体的记录。我们研究这个依赖更新的证据选择问题,并引入契约记忆编译器(CMC)。在看到问题之前,CMC使用语言模型识别历史中的关系并记录每个关系陈述的位置。它应用后续更新来确定当前关系,从问题中命名的实体出发遵循这些关系,并在一次调用中将相应的原始记录传递给答案模型。因此,当前状态决定读取哪些证据,而不是仅仅刷新先前选择的上下文中的值。据我们所知,CMC在FactConsolidation上达到了最先进的多跳准确率,总体达到78.25%,在262K时达到61.0%。在提取的关系和答案模型保持不变的情况下,在解析更新之前选择证据会将多跳准确率降低到21.50%。我们还引入了MQuAKE-MemStream,这是一个从MQuAKE-Remastered反事实案例构建的有序记忆流派生数据集。
英文摘要
External memory lets language-model agents answer questions about histories too long for the answer model's context window. Updates create a harder problem than retrieving a recent fact: changing one relation can redirect a multi-hop question to records about an entity absent from the question. We study this update-dependent evidence selection problem and introduce the Contract Memory Compiler (CMC). Before seeing a question, CMC uses a language model to identify relations in the history and record where each one was stated. It applies later updates to determine the current relations, follows them from entities named in the question, and passes the corresponding original records to the answer model in one call. Thus the current state determines which evidence is read, rather than merely refreshing values in a previously selected context. To the best of our knowledge, CMC achieves state-of-the-art multi-hop accuracy on FactConsolidation, reaching 78.25% overall and 61.0% at 262K. With the extracted relations and answer model held fixed, selecting evidence before resolving updates reduces multi-hop accuracy to 21.50%. We also introduce MQuAKE-MemStream, a derived dataset of ordered memory streams built from MQuAKE-Remastered counterfactual cases.