arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

提示空间元学习无法跨用户迁移:基于冻结大语言模型的负面结果

Prompt-Space Meta-Learning Does Not Transfer Across Users: A Frozen-LLM Negative Result

Liam Byrne, David Dylan, Orla Fitzgerald, Eoin Doyle, Ciara Nolan, Padraig Lynch, Sinead Gallagher

arXiv 2609.01615首次发表:更新:

发表机构

Trinity College Dublin; University College Dublin; Dublin City University(都柏林圣三一大学; 都柏林大学学院; 都柏林城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对冻结LLM的提示空间元学习,通过Muse方法在两个个性化基准上证实其无法跨用户迁移,归因于元目标崩溃,提出可分离混淆因素的对照协议。

AI 中文摘要

将冻结大语言模型(LLM)针对单个用户进行个性化定制,常被表述为提示空间中的元学习问题:每个用户对应一个任务,目标是找到一个共享的自然语言适配策略,在给定少量该用户的标记交互样本时,为该用户配置冻结模型。该表述颇具吸引力,因其与主干无关且可复用提示优化机制,但该领域极少验证优化后的元目标是否编码了可跨用户迁移的适配能力,而非通用指令质量。本文通过Muse(基于共享进化的用户适配元学习)研究该问题:Muse通过反射式提示进化,在元训练用户群体上进化出单个共享适配提示,将其冻结后零样本应用于保留用户;匹配对照组隔离了措辞和选择混淆因素带来的学习效果。在两个标准个性化基准(LaMP-2分类任务、LaMP-3评分任务)上,各含200个保留用户,Muse未显著优于自身未进化的种子提示,也未优于在不匹配用户-支持对上进行元训练的结构破坏对照组,且在评分任务中被普通少样本检索方法超越(平均绝对误差差值+0.175,p<0.001)。本文将这些结果归因于单一机制:元目标崩溃——元验证目标对用户-支持对应关系是否真实具有统计不变性(LaMP-2上p=0.555,LaMP-3上p=0.622),因此无法被优化为可迁移的适配能力,反而奖励指令打磨和验证过拟合。种子提示、错误支持和不变性预言机对照组形成了可复用的协议,可将学习到的适配能力与这些混淆因素分离。

英文摘要

Personalizing a frozen large language model (LLM) to individual users is often framed as a meta-learning problem in prompt space: each user is a task, and one seeks a shared natural-language adaptation policy that, given a handful of the user's labeled interactions, configures the frozen model for that user. The framing is attractive because it is backbone-agnostic and reuses the machinery of prompt optimization, yet the field rarely tests whether the optimized meta-objective encodes transferable cross-user adaptation rather than generic instruction quality. We study this question with Muse (Meta-learned User-adaptation via Shared Evolution), which evolves a single shared adaptation prompt over a meta-train user population by reflective prompt evolution, freezes it, and applies it zero-shot to held-out users; matched controls isolate learning from confounds of phrasing and selection. On two standard personalization benchmarks (LaMP-2 categorization and LaMP-3 rating) over 200 held-out users each, Muse does not significantly improve on its own un-evolved seed prompt or on a structure-broken control that meta-trains on mismatched user-support pairs, and is dominated by plain few-shot retrieval on the rating task (Delta MAE +0.175, p < 0.001). We attribute these outcomes to a single mechanism, meta-objective collapse: the meta-validation objective is statistically invariant to whether the user-support correspondence is genuine (p=0.555 on LaMP-2, p=0.622 on LaMP-3), so it cannot be optimized into transferable adaptation and instead rewards instruction polish and validation overfitting. The seed-prompt, wrong-support, and invariance-oracle controls form a reusable protocol that separates learned adaptation from these confounds.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑