发表机构
University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过访谈提炼八个主题,提出 Reconsider 程序(五项检查、四种处理模式),在 80 个场景上评估,两位 LLM 评审员以正净边际支持其用于 AI 伴侣的记忆使用决策。
AI 中文摘要
记忆可以维持 AI 伴侣关系,然而即使回忆准确,使用它也可能不合时宜。两轮形成性访谈(n = 6 探索性,n = 8 聚焦记忆)促使我们思考:伴侣在使用过去的信息之前应考虑什么。八个主题催生了 Reconsider,一个包含五项检查和四种处理模式的单次调用程序,在跨越五个模型的 80 个场景上进行了评估,共 400 个盲法模型内配对。两位 LLM 评审员以 +15 和 +23 个百分点的净边际支持 Reconsider,其中三个模型的引导区间排除零,但 GPT 和 Claude 除外。评估者分析将评审员评分差异与模型家族联系起来,一项初步的匹配指导对照(隔离记忆特定内容)给出了正边际。我们贡献了一个基于访谈的记忆使用设计框架,以及一项审视其自身评估者的评估。
英文摘要
Memory can sustain AI companionship, yet even accurate recollection can be inappropriate to use. Two rounds of formative interviews with 14 users (n = 6 exploratory, n = 8 memory-focused) motivate asking what a companion should consider before using past information. Eight themes inform Reconsider, a single-call procedure with five checks and four handling modes, evaluated on 80 scenarios across five models over 400 blinded within-model pairs. Two LLM judges favored Reconsider by net margins of +15 and +23 percentage points, with bootstrap intervals excluding zero for three of five models but not for GPT or Claude. Evaluator analysis linked judge scoring differences to model family, and a preliminary matched-guidance control isolating memory-specific content gave positive margins. We contribute an interview-grounded design framework for memory use and an evaluation that scrutinizes its own evaluators.
Comments17 pages, 1 figure, 3 tables. Zihan Guo and Roxy He contributed equally to this work and share first authorship