发表机构
Yonsei University; Hyundai Motors Company; Korea Institute of Science and Technology(延世大学; 现代汽车公司; 韩国科学技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出 CCS 方法,通过轻量适配器动态控制语言模型解码时的上下文影响,在单数据集训练后实现领域内及四个分布外基准的个性化提升,且降低推理成本。
AI 中文摘要
对语言模型(LMs)进行个性化以适配单个用户偏好,是让模型响应符合多样目标与背景的关键。现有方法通常为每个用户训练单独的适配器,或学习依赖于用户的奖励模型。尽管这些方法明确针对每个用户优化,但它们必须从有限观测中学习,因此存在数据稀疏性,且对未见用户和领域的泛化能力较差。上下文学习(ICL)与上下文引导(CoS)可通过直接基于用户上下文调节基础语言模型、利用其预训练能力,无需针对每个用户训练,从而实现更有效的个性化。然而,二者均未在解码步骤中调节上下文的影响:ICL 对上下文影响不加控制,而 CoS 应用固定引导系数,且每一步需要两次语言模型前向传播。我们提出谨慎上下文引导(CCS),它在冻结的骨干语言模型上添加轻量适配器,以在每个 token 处决定用户上下文是否影响生成以及影响强度。该适配器从基于上下文的专家语言模型学习此行为,并在上下文无帮助时保留基础语言模型。仅在一个数据集上训练的单个 CCS 适配器,在领域内以及四个分布外的个性化基准上均提升了生成质量,展现出对新用户和领域的稳健泛化能力。CCS 还避免了针对每个用户的微调,以及 CoS 所需的额外基于上下文的前向传播,大幅降低了推理成本。
英文摘要
Personalizing language models (LMs) to individual user preferences is essential for aligning responses with diverse goals and backgrounds. Existing methods typically train a separate adapter for each user or learn a reward model whose scores depend on the user. Despite explicitly optimizing for each user, these methods must learn from limited observations and therefore suffer from data sparsity and poor generalization to unseen users and domains. In-context learning (ICL) and Context Steering (CoS) can instead provide more effective personalization by conditioning the base LM directly on user context and leveraging its pretrained capabilities without per-user training. Yet neither adapts the influence of that context across decoding steps: ICL leaves it uncontrolled, whereas CoS applies a fixed steering coefficient and requires two LM forward passes per step. We propose Cautious Context Steering (CCS), which adds a lightweight adapter to a frozen backbone LM to decide at each token whether and how strongly user context should affect generation. The adapter learns this behavior from an oracle context-conditioned LM and preserves the base LM when the context is not helpful. A single CCS adapter trained on only one dataset improves generation quality both in-domain and across four out-of-distribution personalization benchmarks, demonstrating robust generalization to new users and domains. CCS also avoids per-user fine-tuning and the additional context-conditioned forward pass required by CoS, substantially reducing inference cost.
Comments11 pages, 3 figures