arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26853cs.LGcs.AI

COPE:基于用户嵌入与自我评估的稀疏用户反馈下LLM持续个性化

COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation

Ruike Cao, Fugen Yao, Liang Dong, Jian Xu, Guanjun Jiang, Li Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

COPE通过为每位用户分配可学习嵌入并结合自我评估生成代理奖励,在稀疏反馈下实现LLM的持续个性化优化,实验证明其优于现有基线并与检索增强提示互补。

中文摘要 AI 辅助

尽管大型语言模型(LLMs)在各种基准测试中取得了显著成果,但其与规范性价值观的对齐往往导致回应同质化,无法满足多样化的用户偏好。现有的免训练方法通常通过提示工程占用宝贵的上下文窗口,而基于训练的方法通常在训练后保持静态,无法支持现实场景中所需的持续优化。为解决这些挑战,我们提出了COPE(基于个性化嵌入与自我评估的持续优化),一种专为现实世界激励的稀疏用户反馈交互场景设计的新型优化框架。我们的框架为每位用户分配可学习的个性化嵌入,并在单次更新步骤中协同整合偏好捕获、自我评估校准和个性化响应优化。我们方法的一个关键创新是利用自我评估生成代理奖励,使得即使在缺乏显式用户反馈的情况下也能实现模型的持续更新。实验表明,COPE在稀疏反馈下持续优于强力的免训练和基于训练的基线方法,并且与检索增强提示(RAP)保持互补。进一步的分析证实了COPE可靠的自我评估、有意义的偏好模式、稳定的通用能力,以及在偏好变化和替代评估器下的鲁棒性。

英文摘要

While Large Language Models (LLMs) have achieved remarkable results across various benchmarks, their alignment with normative values often results in homogenized responses that fail to address diverse user preferences. Existing training-free methods often occupy valuable context windows through prompt engineering, while training-based methods typically remain static post-training, failing to support the continual optimization required in real-world settings. To address these challenges, we propose COPE (Continual Optimization with Personalized embedding and self-Evaluation), a novel optimization framework tailored for real-world-motivated interaction settings with sparse user feedback. Our framework assigns learnable personalized embeddings to each user and synergistically integrates preference capture, self-evaluation calibration, and personalized response optimization within a single update step. A key innovation of our method is the use of self-evaluation to generate proxy rewards, enabling continuous model updates even when explicit user feedback is unavailable. Experiments show that COPE consistently outperforms strong training-free and training-based baselines under sparse feedback, and remains complementary to Retrieval-Augmented Prompting (RAP). Further analyses confirm COPE's reliable self-evaluation, meaningful preference patterns, stable general capabilities, and robustness under shifting preferences and alternative evaluators.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Qwen Business Unit of Alibaba(阿里巴巴千问事业部)

机构由 AI 辅助整理,请以论文原文为准。

↑