假设引导的自蒸馏用于持续个性化
Hypotheses-Guided Self Distillation for Continual Personalization
浏览论文内容
中文总结 AI 辅助
本文提出HypReflect框架,通过推断并优化用户偏好假设、结合假设引导的自蒸馏实现持续个性化,实验显示其在多场景下性能优于基线方法,具备强泛化性与稳定性。
中文摘要 AI 辅助
随着人们在日常生活中与大语言模型(LLM)助手的交互日益频繁,持续适配个人偏好已成为实现长期有效交互的关键。然而,用户偏好很少被完整表述,而是通过异构、潜在且含噪的信号呈现;现有方法依赖原始交互历史或成本高昂的基于奖励的优化来管理个性化。本文提出HypReflect,一种可靠、可扩展的持续个性化框架,该框架从多样的用户信号中推断明确的、感知不确定性的偏好假设,随新证据的积累对这些假设进行反思式优化,并通过假设引导的自蒸馏整合得到的用户模型。在三种个性化场景(在线个性化、多会话交互、隐式行为信号)下开展的实验表明,HypReflect的性能优于包括原始历史方法和增量更新方法在内的一系列基线方法。我们进一步验证了其对未见用户和跨域场景的强泛化能力,以及在上下文预算、可复用假设和更聚焦的个性化方面的稳定性。这些结果表明,通过明确、可修正的用户偏好假设,我们朝着实现可靠且可扩展的持续个性化迈出了一步。
英文摘要
As people increasingly interact with LLM assistants in daily life, continually adapting to individual preferences has become essential for effective long-term interactions. However, user preferences are rarely stated in full, and instead emerge through heterogeneous, latent, and noisy signals, with existing methods relying on raw interaction histories or costly reward-based optimization to manage personalization. We introduce HypReflect, a reliable, scalable framework for continual personalization that infers explicit, uncertainty-aware preference hypotheses from diverse user signals, reflectively refines them as new evidence accumulates, and incorporates the resulting user model through hypotheses-guided self-distillation. Experiments across three personalization settings: online personalization, multi-session interactions, and implicit behavioral signals, show that HypReflect outperforms a range of baselines, including raw-history and incremental-update methods. We further demonstrate strong generalization to unseen users and cross-domain settings, along with stability across context budgets, reusable hypotheses, and more focused personalization. These results suggest a step towards reliable and scalable continual personalization through explicit, revisable user preference hypotheses.
发表机构
- University of British Columbia(不列颠哥伦比亚大学)
- Megagon Labs(美伽贡实验室)
机构由 AI 辅助整理,请以论文原文为准。