UpgradeBench:面向微调LLM专用模型升级的决策导向基准
UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists
浏览论文内容
中文总结 AI 辅助
该研究提出UpgradeBench基准,针对LLM专用模型升级的决策问题,在多模型序列与任务上验证了适配器迁移等策略的效果,实现低成本高效升级。
中文摘要 AI 辅助
企业会为开放权重语言模型维护特定任务的适配器,每次新基础模型发布都需做出迁移决策:保留现有专用模型、迁移适配器、从保留行为中刷新或重新训练。现有迁移研究仅评估孤立模型对,未在真实模型发布序列中研究这些选择。我们提出UpgradeBench,这是一个决策驱动的纵向基准,涵盖4次连续Qwen版本、1个续 checkpoint、6个任务和2种模型规模,并补充了具有已知训练谱系的OLMo checkpoint。该基准拆解了三个核心问题:新checkpoint是否会改进固定配方重新训练的专用模型性能、专用资产是否可跨版本迁移、可用的恢复资源是什么。我们观察到升级收益在任务-规模-发布场景中存在差异:部分重新训练基线有所提升,其余则处于训练噪声范围内;耐用性从文本到SQL的不到1个发布间隔,到意图分类的超过14个月不等。直接适配器迁移不依赖架构或模型家族:在OLMo上,保留率在46B token的持续预训练时为0.88-0.99,在2.9T token时降至0;退火和模型拼接未引入额外损害,可移植性随持续预训练距离衰减。在保留输入数据的情况下,教师重新标注可恢复目标-基础专用模型,无需新的黄金标注,但计算节省无法保证。对33个升级场景模拟固定决策策略,产生0.37个百分点的平均质量遗憾,且无行为退化,计算和标签成本仅为完全重新训练的三分之一。基于256个提示的轻量级CKA探测可预测跨版本适配器可移植性(8个模型对的斯皮尔曼相关系数为0.74)。我们发布了逐示例预测、成本日志、拆分清单和评估代码。
英文摘要
Organizations maintain task-specific adapters for open-weight language models, and each new base-model release forces a migration decision: retain existing specialists, port adapters, refresh from preserved behavior, or retrain. Prior transfer work evaluates isolated model pairs, without studying these choices across real model release sequences. We present UpgradeBench, a decision-driven longitudinal benchmark covering four consecutive Qwen releases, one continuation checkpoint, six tasks, and two model sizes, augmented by OLMo checkpoints with known training lineage. The benchmark disentangles three core questions: whether a new checkpoint improves fixed-recipe retrained specialist performance, whether specialization assets transfer across versions, and what recovery resources are usable. We observe upgrade gains differ across task-scale-release episodes: some retrained baselines improve while others stay within training noise, with durability ranging from under one release interval for text-to-SQL to over fourteen months for intent classification. Direct adapter copying depends neither on architecture nor model family: on OLMo, retention drops from 0.88-0.99 at 46B-token continued pretraining to zero at 2.9T tokens; annealing and model souping introduce no extra harm, with portability decaying with continued-pretraining distance. Given preserved input data, teacher relabeling recovers target-base specialists without fresh gold annotations, though compute savings are not guaranteed. Simulating a fixed decision policy over 33 upgrade episodes yields 0.37pp mean quality regret with zero behavioral regressions at one-third the compute and label cost of full retraining. A lightweight CKA probe over 256 prompts predicts cross-version adapter portability (Spearman 0.74 across eight model pairs). We release per-example predictions, cost logs, split manifests, and evaluation code.
发表机构
- Alibaba Group(阿里巴巴集团)
- Cheung Kong Graduate School of Business(长江商学院)
机构由 AI 辅助整理,请以论文原文为准。