发表机构
Xiaohongshu Inc.; National University of Singapore; University of Science and Technology of China; Peking University(小红书公司; 新加坡国立大学; 中国科学技术大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GRASP通过评分标准感知的在线策略自蒸馏,将用户特定评分标准转化为令牌级监督,并引入教师验证,在LaMP-QA上实现最先进性能。
AI 中文摘要
大语言模型个性化旨在生成符合个体用户偏好和需求的响应。用户特定的评分标准使这些期望明确化,为满意的答案应涵盖的内容提供直接监督。然而,现有的评分标准引导方法仅在粗粒度上利用这种引导,要么使用评分标准来监督相关方面的预测以供后续生成,要么将方面覆盖减少为强化学习的单一响应级奖励。这在指定个性化答案应包含的内容与教导模型如何生成它之间留下了差距。为弥合这一差距,我们提出GRASP,一种用于大语言模型个性化的评分标准感知的在线策略自蒸馏框架,将用户特定的评分标准方面转化为细粒度的、令牌级监督。具体来说,GRASP将无评分标准的学生与有评分标准的教师配对,后者额外接收目标用户特定的评分标准。通过沿学生生成的在线策略轨迹对齐它们的下一令牌分布,GRASP将教师的评分标准条件引导转移给学生,将用户特定的语义要求转化为密集的令牌级监督。由于有评分标准的教师仍可能产生不充分的监督,我们进一步引入基于评分标准的教师验证(RTV),仅保留教师充分覆盖目标方面的实例,从而提高监督质量和训练效率。在LaMP-QA基准上进行个性化问答的实验表明,GRASP在多个骨干模型上实现了最先进的性能,支持了评分标准引导的令牌级监督对个性化的有效性。为确保可复现性,我们的代码可在以下网址获取:https://this https URL。
英文摘要
LLM personalization aims to generate responses aligned with individual users' preferences and needs. User-specific rubrics make these expectations explicit, providing direct supervision on what a satisfactory answer should cover. Existing rubric-guided approaches, however, exploit such guidance only at a coarse granularity, either by using rubrics to supervise the prediction of relevant aspects for subsequent generation or by reducing aspect coverage to a single response-level reward for reinforcement learning. This leaves a gap between specifying what a personalized answer should contain and teaching the model how to generate it. To bridge this gap, we propose GRASP, a rubric-aware on-policy self-distillation framework for LLM personalization that turns user-specific rubric aspects into fine-grained, token-level supervision. Specifically, GRASP pairs a rubric-free student with a rubric-informed teacher that additionally receives the target user-specific rubrics. By aligning their next-token distributions along on-policy trajectories generated by the student, GRASP transfers the teacher's rubric-conditioned guidance into the student, translating user-specific semantic requirements into dense token-level supervision. Since rubric-informed teachers can still produce inadequate supervision, we further introduce Rubric-based Teacher Validation (RTV), which retains only instances where the teacher sufficiently covers the target aspects, improving both supervision quality and training efficiency. Experiments on the LaMP-QA benchmark for personalized question answering demonstrate that GRASP achieves state-of-the-art performance across multiple backbones, supporting the effectiveness of rubric-guided token-level supervision for personalization. To ensure reproducibility, our code is available at https://github.com/SnowCharmQ/GRASP.