VISA: 通过屏蔽适应进行价值注入以实现个性化大语言模型对齐
VISA: Value Injection via Shielded Adaptation for Personalized LLM Alignment
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
VISA通过屏蔽适应机制,有效缓解大语言模型对齐税,实现精细化价值观对齐与语义完整性保持。
AI中文摘要:
对齐大型语言模型(LLMs)与细腻的人类价值观仍然是一个关键挑战,因为现有的方法如人类反馈强化学习(RLHF)通常只能处理粗粒度属性。在实践中,通过任务特定数据集对LLMs进行微调以优化价值对齐不可避免地会带来对齐税:模型的预校准价值观系统由于训练数据中的潜在偏见吸收而显著漂移,同时微调过程还会导致生成响应中的严重幻觉和语义信息丢失。为了解决这个问题,我们提出了VISA(通过屏蔽适应进行价值注入),一种闭环框架,旨在导航这种权衡。VISA的架构特征包括高精度价值观检测器、语义到价值观翻译器和核心价值观重写器。价值观重写器通过组相对策略优化(GRPO)进行训练,使用一个复合奖励函数同时优化细粒度价值观精度和语义完整性保持。通过学习一个最优策略来平衡这些竞争性目标,VISA有效减轻了对齐税,同时保持对原始知识的忠诚。我们的实验表明,这种方法能够在保持模型事实一致性和通用能力的同时,对模型的价值表达进行精确控制,显著优于标准微调方法和基于提示的基线方法,包括GPT-4o。
英文摘要:
Aligning Large Language Models (LLMs) with nuanced human values remains a critical challenge, as existing methods like Reinforcement Learning from Human Feedback (RLHF) often handle only coarse-grained attributes. In practice, fine-tuning LLMs on task-specific datasets to optimize value alignment inevitably incurs an alignment tax: the model's pre-calibrated value system drifts significantly due to latent bias absorption from training data, while the fine-tuning process also causes severe hallucinations and semantic information loss in generated responses. To address this, we propose VISA (Value Injection via Shielded Adaptation), a closed-loop framework designed to navigate this trade-off. VISA's architecture features a high-precision value detector, a semantic-to-value translator, and a core value-rewriter. The value-rewriter is trained via Group Relative Policy Optimization (GRPO) with a composite reward function that simultaneously optimizes for fine-grained value precision, and the preservation of semantic integrity. By learning an optimal policy to balance these competing objectives, VISA effectively mitigates the alignment tax while staying loyal to the original knowledge. Our experiments demonstrate that this approach enables precise control over a model's value expression while maintaining its factual consistency and general capabilities, significantly outperforming both standard fine-tuning methods and prompting-based baselines, including GPT-4o.