发表机构
The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLM个性化中记忆利用的问题,提出解耦评估范式,发现智能体常能回忆用户偏好却未在行为中体现,健康相关偏好的利用薄弱且风险最高。
AI 中文摘要
随着大语言模型(LLM)智能体发展为个性化伙伴,记忆已成为核心能力。然而,LLM面临知识利用问题:即使相关用户偏好完全存在于上下文,智能体也可能无法基于该偏好行动。当智能体在应考虑先前分享的用户偏好的情境中未能调整回应时,尚不明确是模型未能记住该信息,还是记住了但未使用它。为分离这种失效,我们引入一种解耦评估范式,对同一用户偏好实施配对的“了解(Know)”测试与“执行(Act)”测试。我们在16个系统和5种记忆架构上开展大规模实验,评估嵌入3种表达强度级别的1000个偏好。结果显示,“了解”与“执行”结果间存在巨大差距:智能体常通过用户偏好的回忆测试,但未在配对行为场景中体现同一偏好。尽管记忆架构能缩小该差距,但与健康和治疗相关的偏好利用仍尤其薄弱,这类偏好的执行失败会带来最大的现实风险。
英文摘要
As large language model (LLM) agents evolve into personalized companions, memory has emerged as a core capability. However, LLMs face a knowledge utilization problem: they may fail to act on relevant user preferences even when they are fully present in context. When an agent fails to tailor its response in a context where previously shared user preferences should matter, it is unclear whether the model failed to remember that information or remembered it but failed to use it. To isolate this breakdown, we introduce a decoupled evaluation paradigm that administers paired Know and Act tests to the same user preference. We conduct large-scale experiments across 16 systems and five memory architectures, evaluating 1,000 preferences embedded at three levels of expression strength. Our results show a large gap between Know and Act outcomes: agents often pass the recall test for a user preference but fail to reflect that same preference in the paired behavioral scenario. While memory architectures reduce this gap, utilization remains especially weak for health and therapy-related preferences, where failures to act carry the greatest real-world stakes.