发表机构
AAII, University of Technology Sydney(悉尼科技大学AAII)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对记忆增强助手难以判断偏好何时适用的问题,提出PairPref基准,通过仅改变情境的1,227对样本测试模型,发现多数模型选择得分尚可但自由生成中仅3.6%-18.3%正确,表明模型仍缺乏情境化偏好使用能力。
AI 中文摘要
记忆增强型助手使用检索到的偏好来引导其回答。情境中的微小变化可能改变某个偏好是否适用,却几乎不影响其检索相似度。现有的记忆基准通常测试系统能否存储和检索偏好,而较少关注这些偏好何时应当被应用。我们引入了PairPref,一个关于情境化偏好使用的基准。每一对样本仅改变情境,保持偏好、请求和四个候选回答不变。该偏好在这两种情境中均保持有效。在选择轨中,模型必须选出仅在适当处应用偏好的回答。在自由生成轨中,模型必须在未看到候选回答的情况下决定何时应用该偏好。两个轨均使用相同的1,227对样本,涵盖45个偏好和八个情境类别。我们评估了八个模型,其中大多数模型的选择得分(Δ)在51到65分之间。然而,在自由生成中,两种回答均分别适用于各自情境的样本仅占3.6%至18.3%。即使在检索记忆更少、呈现格式不同以及提示更严格的情况下,模型仍会在两种情境中继续应用该偏好。这些结果表明,模型仍然难以判断用户偏好何时适用并据此作出回应。
英文摘要
Memory-augmented assistants use retrieved preferences to guide their responses. A small change in the situation can change whether a preference is appropriate while barely affecting its retrieval similarity. Memory benchmarks typically test whether systems store and retrieve preferences, with less attention to when those preferences should apply. We introduce PairPref, a benchmark of contextual preference use. Each pair changes only the situation, keeping the preference, request, and four candidate replies fixed. The preference remains valid in both situations. In the selection track, models must choose the reply that applies the preference only where appropriate. In the free-generation track, they must decide when to apply it without seeing candidate replies. Both tracks use the same 1,227 pairs across 45 preferences and eight situation categories. We evaluate eight models, most of which achieve selection scores ($Δ$) of 51 to 65 points. In free generation, however, both responses are appropriate for their respective situations in only 3.6\% to 18.3\% of pairs. Models continue to apply the preference in both situations even with fewer retrieved memories, alternative presentation formats, and a stricter prompt. These results show that models still struggle to judge when user preferences apply and respond accordingly.