arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在回答之前:大小匹配记忆构建下的证据充分性

Before Answering: Evidence Sufficiency under Size-Matched Memory Construction

Joyanta Jyoti Mondal, Md. Shifatul Ahsan Apurba, Mridul Banik, Md Masud Al Mahmud, Ibne Farabi Shihab

arXiv 2609.32269首次发表:更新:

AI 中文总结

本研究提出大小匹配的记忆构造方法消除证据不足任务中的标签捷径,并评估MemSafe估计器,在多个多跳问答数据集上显著优于词汇基线,同时揭示构造方式对检测性能的关键影响。

AI 中文摘要

从压缩或检索记忆中回答问题的智能体必须识别查询所需的证据是否已不在记忆中。针对此任务的基准通常通过删除支持段落来构造证据不足的示例。我们表明,这种构造通过记忆大小泄露了标签:在MuSiQue上,仅统计段落数量的分类器在检测不安全记忆时达到了0.979的ROC曲线下面积(AUROC),高于我们最初评估的词汇估计器。我们提出了一种大小匹配的构造方法,该方法可证明地消除了这一捷径,并使用它来研究MemSafe,一种将查询与每个记忆单元交叉编码并通过集合变换器聚合单元的估计器。在三个多跳问答数据集和五个随机种子上,MemSafe在MuSiQue和HotpotQA上分别达到0.968和0.983的AUROC,比词汇基线高出0.26到0.39,而第三个数据集2WikiMultiHopQA已饱和。一个带有逻辑回归头的冻结预训练交叉编码器已经缩小了MuSiQue上词汇基线与MemSafe之间差距的41%。同时,MemSafe在MuSiQue发布的不可能回答问题上比弱基线退化更多,在SQuAD 2.0上仅达到0.639的AUROC,并且需要数千个临床训练示例才能优于基于特征的估计器。作为7B阅读器的门控,它在5%覆盖率下将已回答问题的错误率从0.850降至0.631,优于阅读器置信度,并且在平均情况下优于真实完整性标签,尽管7B LLM评判器在10%覆盖率下是更好的门控。这些结果表明,证据不足的构造方式与检测它的估计器同样重要。

英文摘要

Agents that answer questions from compressed or retrieved memory must recognize when the evidence a query needs is no longer in memory. Benchmarks for this task usually create insufficient-evidence examples by deleting supporting passages. We show that this construction leaks the label through memory size: on MuSiQue, a classifier that only counts paragraphs reaches an area under the ROC curve (AUROC) of $0.979$ for detecting unsafe memory, higher than the lexical estimator we initially evaluated. We propose a size-matched construction that provably removes this shortcut, and use it to study MemSafe, an estimator that cross-encodes the query with each memory unit and aggregates the units with a set transformer. Across three multi-hop question answering datasets and five seeds, MemSafe reaches $0.968$ and $0.983$ AUROC on MuSiQue and HotpotQA, $0.26$ to $0.39$ above a lexical baseline, while the third dataset, 2WikiMultiHopQA, is saturated. A frozen pretrained cross-encoder with a logistic head already closes $41\%$ of the MuSiQue gap between the lexical baseline and MemSafe. At the same time, MemSafe degrades more than a weak baseline on the unanswerable questions released with MuSiQue, reaches only $0.639$ AUROC on SQuAD~2.0, and needs several thousand clinical training examples before it outperforms a feature-based estimator. Used as a gate for a 7B reader, it reduces the error rate on answered questions from $0.850$ to $0.631$ at $5\%$ coverage, outperforming both reader confidence and, on average, the ground-truth integrity label, although a 7B LLM judge is the better gate at $10\%$ coverage. These results indicate that the way insufficient evidence is constructed matters as much as the estimator that detects it.

Comments24 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑