基于激活传播的低资源大语言模型偏好适配
Low-Resource Preference Adaptation of LLMs via Activation-Based Label Propagation
浏览论文内容
中文总结 AI 辅助
针对低资源场景下LLMs偏好优化受标注成本限制的问题,本文提出基于少量标注数据训练的轻量线性探针,利用激活结构标注大量未标注数据,其性能优于直接训练且接近使用更多标注数据的基线。
中文摘要 AI 辅助
将大语言模型(LLMs)适配至用户特定偏好常受限于人工标注成本,在低资源场景中偏好无法由LLMs自身可靠标注(如因文化、主观或个性化语境),偏好优化难以实施。本文研究语言模型如何在中间表示中编码偏好信息,发现所选与拒绝响应的激活在各层形成 distinct 聚类,即便在预训练模型中亦是如此;值得注意的是,该结构经标准数据集对齐后会增强,但当目标偏好与模型对齐的偏好不同时会被抹去,表明对齐后的LLMs对非主流群体判断能力差。利用该结构,本文提出在少量标注偏好对(≤500)上训练轻量线性探针,并用其为下游偏好优化标注大型未标注数据集(50K+)。本文在不同数据集、偏好优化方法及模型规模上系统评估该方法,发现相同标注预算下,该方法始终优于直接训练,且在多数场景中,与使用50-100倍更多标注数据训练的基线方法相比仍具竞争力。代码可在该https URL获取。
英文摘要
Adapting large language models to user-specific preferences is often constrained by the cost of human annotation, making preference optimisation impractical in low-resource settings where preferences cannot be reliably labelled by LLMs themselves, e.g., due to cultural, subjective, or personalised contexts. In this paper, we investigate how language models encode preference information in their intermediate representations, finding that activations from chosen and rejected responses form distinct clusters across layers, even in pretrained models. Strikingly, this structure is strengthened by alignment on canonical datasets but erased when the target preferences differ from those the model was aligned on, suggesting aligned LLMs are poor judges for non-mainstream populations. Exploiting this structure, we propose training a lightweight linear probe on a few labelled preference pairs ($\leq$500) and using it to annotate large unlabelled datasets (50K+) for downstream preference optimisation. We systematically evaluate this approach across different datasets, preference optimisation methods and model scales and find that our method consistently outperforms direct training given the same annotation budget, and remains competitive against baselines trained on $50-100\times$ more labelled data in the majority of our settings. Code is available at https://github.com/alessioGalatolo/activ-pref-probe.
发表机构
- Uppsala University(乌普萨拉大学)
机构由 AI 辅助整理,请以论文原文为准。