arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MIRROR:从模仿到内化——LLM个性化中的转变

MIRROR: From Imitation to Internalization in LLM Personalization

Huayi Lai, Jicheng Yang, Min Yi, Chong Meng

arXiv 2610.09795首次发表:更新:

发表机构

Baidu Inc.(百度公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MIRROR框架,通过参考揭示的在线策略自蒸馏将LLM个性化从模仿转向偏好内化,并引入MIRROR-F插件,在多个基准上取得领先性能并减少灾难性遗忘。

AI 中文摘要

对个性化大语言模型(LLM)的需求正从风格模仿转向内容质量。我们研究了在现有微调范式中,自蒸馏(self-distillation)能否弥合这一差距。为解决这一局限,我们提出了MIRROR(通过内化参考揭示的在线策略反思实现元个性化),一种新颖的自蒸馏框架,将LLM个性化从模仿转向偏好内化。首先,我们用参考揭示的在线策略自蒸馏取代参考标记模仿,使模型沿自身生成轨迹的下一词分布与其参考条件化自身的分布对齐,从而内化用户偏好而非复现参考文本。其次,我们引入MIRROR-F,一种焦点插件,通过选择性监督信息性参考标记来增强在线策略分布对齐,从而在保留用户特定表达的同时强化内容生成。在三个个性化生成基准、两种模型规模以及互补的基于参考和基于LLM的评估中,MIRROR和MIRROR-F取得了领先的整体个性化性能和优越的文本质量,同时在三个未见过的个性化生成任务上表现出比基于SFT的基线更少的灾难性遗忘。这些改进在模型规模和应用程序场景中保持一致,转化为LLM个性化任务中性能的提升。

英文摘要

The demand for personalized LLMs is shifting from style imitation toward content quality. We investigate whether self-distillation can bridge this gap in existing fine-tuning paradigm. To address this limitation, we introduce MIRROR(Meta- personalization by Internalizing Reference-Revealed On-policy Reflections), a novel self-distillation framework that shifts LLM personalization from imitation toward preference internalization. First, we replace reference-token imitation with reference-revealed on-policy self-distillation, aligning the model's next-token distributions along its own generation trajectories with those of its reference-conditioned self, thereby internalizing user preferences rather than reproducing reference wording.Second, we introduce MIRROR-F, a focal plug-in that augments on-policy distributional alignment with selective supervision over informative reference tokens, thereby strengthening content generation while preserving user-specific expression. Across three personalized generation benchmarks, two model scales, and complementary reference-based and LLM-based evaluations, MIRROR and MIRROR-F achieve leading overall personalization performance and superior text quality, while exhibiting less catastrophic forgetting than SFT-based baselines on three unseen personalized generation tasks. The gains are consistent across model scales and application scenarios, translating to improved performance in LLM personalization tasks.

Comments36 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑