arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过语言强化学习实现大语言模型个性化的偏好适配学习

Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning

Yuting Liu, Wei Wu, Jianzhe Zhao, Guibing Guo

arXiv 2608.09507首次发表:更新:

发表机构

Software College, Northeastern University; Ant International(东北大学软件学院; 蚂蚁国际)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对LLM个性化中通用偏好摘要冗余问题,提出无训练元学习框架AlignXada,经语言强化学习优化后在多任务多模型上提升性能且适配效果优于RAG。

AI 中文摘要

自然语言用户偏好为大语言模型(LLM)个性化提供了可解释的接口,但通用偏好摘要常包含与特定下游任务无关的信息。直接提供完整偏好摘要会浪费上下文容量并引入跨任务干扰,而手动设计特定任务的偏好视图难以规模化。本研究探讨「特定任务偏好适配」:给定通用用户偏好摘要和下游任务,推导保留足够决策相关证据同时移除冗余上下文的任务条件表示。为此,我们提出AlignXada,这是一种无训练元学习框架,可诱导可复用的文本精调策略,将通用偏好摘要适配为特定任务版本。该精调策略由元学习器通过语言强化学习迭代优化。在13个任务和3个下游模型(共39个任务-模型单元)上,AlignXada平均提升3.82个点,改善33个单元,同时仅保留原始概要标记的22.8%,并在36个单元上优于检索增强生成(RAG)。扩展的忠实性分析进一步表明,精调后的概要在很大程度上仍基于源偏好,同时保留了任务相关的个性化信号,这表明概要侧适配可作为终身个性化智能体通用记忆构建的实用补充。

英文摘要

Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplying the full preference summary therefore wastes context capacity and introduces cross-task distraction, while manually designing task-specific preference views is difficult to scale. In this work, we study \emph{task-specific preference adaptation}: given a universal user preference summary and a downstream task, derive a task-conditioned representation that preserves sufficient decision-relevant evidence while removing redundant context. To this end, we propose \textsc{AlignXada}, a training-free meta-learning framework that induces reusable textual refinement policies for adapting universal preference summaries to task-specific ones. The refinement policy is iteratively optimized by a meta learner through verbal reinforcement learning. Across 13 tasks and three downstream models (39 task--model cells), \textsc{AlignXada} achieves an average gain of 3.82 points, improving 33 cells while retaining only 22.8\% of the original profile tokens and outperforming RAG in 36 cells. An extended faithfulness analysis further shows that the refined profiles remain largely grounded in the source preferences while preserving task-relevant personalization signals, suggesting that profile-side adaptation serves as a practical complement to universal memory construction for lifelong personalized agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑