分离记忆与工作流效应在预测个体答案中的作用
Separating Memory and Workflow Effects in Predicting Individual Answers
浏览论文内容
中文总结 AI 辅助
本研究分离个性化语言智能体的记忆选择与使用方式,提出OwnWords方法,通过BM25检索个人原话进行单次回答,在有限上下文下优于书面记忆,但未超越近期截断,揭示了证据构建过程而非逐字措辞的影响。
中文摘要 AI 辅助
个性化语言智能体既选择记住关于一个人的哪些信息,也选择如何使用这些记忆。在预测已知面试问题的未见答案时,我们将这两种选择分开处理。在来自188人的1,768个任务上,由经过验证的面试前缀构建的具体记忆,其得分比特质描述高出0.0158(95%全人区间[0.0044, 0.0271])。将两种记忆与一次性生成和三答案融合交叉组合,融合使具体记忆得分降低0.0123([-0.0189, -0.0056]);提示式选择器和训练式选择器均未明显优于随机候选。对更长的未重写源记录进行一次调用,其得分超过所有记忆条件。在有限上下文预算下,OwnWords使用BM25检索该人的句子,并在一次调用中给出答案。它在基准之外的500人上优于书面记忆(+0.0127, [+0.0037, +0.0217];早期留出测试结果不显著),并在300人上跨四个预算平均优于书面记忆(平均+0.0218, [+0.0138, +0.0298]),后者结果在114人上重复。它未明显优于近期截断。这些结果比较了证据构建过程;它们并未分离逐字措辞的影响。在Twin-2K-500上,OwnWords预测序数调查答案比书面记忆更接近,但未提高精确选择准确率,并在两个样本之一中降低了该准确率。面试评分使用基于模型的、无人工评分的、基于内容的评分标准,且原始基准的参与者在开发期间被看到。这些结果描述了所测试过程的特点,而非一般人类预测上限。
英文摘要
Language agents choose what to remember about a person and how to use that memory. We separate these choices when predicting a person's unseen answer to an interview question. On 1,768 tasks from 188 people, a concrete memory from a verified interview prefix outscores a trait description by 0.0158 (95% interval [0.0044, 0.0271]). Crossing both memories with one-shot generation and three-answer fusion, fusion lowers concrete-memory scores by 0.0123 ([-0.0189, -0.0056]); prompted and trained selectors do not detectably beat a random candidate, and one call on the longer, unrewritten record outscores every memory condition. At matched context budgets, OwnWords, one call on the person's BM25-ranked sentences, outperforms the written memory on 500 people outside the benchmark (+0.0127, [+0.0037, +0.0217]; an earlier held-out test was inconclusive) and across four budgets on 300 people (mean +0.0218, [+0.0138, +0.0298]), but does not detectably outperform recency truncation; these comparisons do not isolate verbatim wording. On a survey benchmark it predicts ordinal answers more closely than the memory but is not more accurate on exact choices (less accurate in one of two screener samples). Interview scores use a model-based content rubric without human ratings, and all benchmark and some confirmation participants were seen during development.
发表机构
- University of Cambridge(剑桥大学)
- Stanford University(斯坦福大学)
- Hong Kong Baptist University(香港浸会大学)
- Cookiy Labs(Cookiy实验室)
机构由 AI 辅助整理,请以论文原文为准。