arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12389cs.AI

通过元LoRA学习适应跨领域偏好以实现大语言模型个性化

Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization

Xuefei Wang, Jun Han, Zixuan Wang, Qingkai Zeng, Xiao Wang, Ruijie Wang, Jianxin Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对跨领域个性化的过拟合与负迁移问题,提出PAC-Bayes正则化的Meta-LoRA,分解个性化先验为用户与领域组件,在多基准任务上取得胜率提升。

中文摘要 AI 辅助

跨领域零样本或少样本个性化旨在仅从少量目标领域交互中生成用户偏好的、适用于未见对话领域的响应。现有适应方法难以在稀疏证据下校准更新幅度,因此出现过拟合;而历史迁移方法常将用户偏好与源领域伪影纠缠,导致个性化先验不可靠并产生负迁移。为校准适应与证据质量的匹配,我们提出PAC-Bayes正则化的Meta-LoRA,其使用元学习得到的LoRA初始化作为适应起点和先验中心,同时根据支持集大小和预测不确定性调整更新强度。这限制了稀疏或模糊证据下的过拟合,且允许在证据充足时进行更强的个性化。仅受控适应无法确定哪些偏好应跨领域迁移或如何表达,因此我们将个性化先验在功能上分解为用户和领域组件,使用人类可读的提示词处理稳定偏好,使用拓扑保持的软标记处理特定领域的隐空间条件。在多个基准和个性化任务上的实验显示,我们的方法相较于强基线取得了一致的提升;在HiCUPID数据集上,我们的方法使跨领域胜率下降较最佳竞争基线降低了47.9%,并在未见用户冷启动下使胜率提升了110.2%。

英文摘要

Cross-domain zero- or few-shot personalization aims to generate user-preferred responses in unseen conversational domains from only a handful of target-domain interactions. Existing adaptation methods struggle to calibrate update magnitude under sparse evidence and thus overfit, whereas history-transfer methods often entangle user preferences with source-domain artifacts, yielding unreliable personalization priors and negative transfer. To calibrate adaptation to evidence quality, we propose PAC-Bayes-regularized Meta-LoRA, which uses a meta-learned LoRA initialization as both the adaptation start and prior center, while adjusting update strength according to support-set size and predictive uncertainty. This limits overfitting under sparse or ambiguous evidence while permitting stronger personalization as evidence grows. Controlled adaptation alone does not determine which preferences should transfer across domains or how they should be expressed. We therefore functionally decompose personalization priors into user and domain components, using a human-readable prompt for stable preferences and topology-preserving soft tokens for domain-specific hidden-space conditioning. Experiments across multiple benchmarks and personalization tasks show consistent gains over strong baselines. On HiCUPID, our method reduces cross-domain win-rate degradation by 47.9% relative to the best competing baseline and improves win rate by 110.2% under unseen-user cold start.

↑