发表机构
Athens University of Economics and Business(雅典经济与商业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究资源受限设备上基础模型的最优持续微调问题,将其建模为约束马尔可夫决策过程,提出基于强化学习的演员评论家算法,实验表明该方法在准确率上优于同预算微调方法,且大幅减少微调步骤。
AI 中文摘要
我们研究在资源受限设备上部署的预训练基础模型的最优持续微调问题。每个时隙有新训练数据,控制器面临微调模型并产生计算成本或不微调而丢弃数据的选择。决策后,根据特定应用性能指标衡量当前模型性能。目标是在有限计算预算下学习确定何时在单个任务(如情感分析)上微调模型的最优策略。将此在线决策问题表述为约束马尔可夫决策过程,系统状态包含模型性能、计算预算和数据分布相关性。转移是随机的,提出基于强化学习的方法(即演员评论家算法)解决。还考虑了可预测微调性能的特殊情况,此时问题变为动态规划问题。在广泛使用的文本分类数据集上用大型预训练模型进行实验表明,我们的方法在准确率方面始终比相同计算预算的微调方法高出4%以上,仅需25%的微调步骤就能达到全参数微调准确率的97%。
英文摘要
We study the problem of optimal continual fine-tuning for a pre-trained Foundation Model deployed at a resource-limited device. At each time slot, a new batch of training data arrives, and the controller is faced with two options: either use the data to fine-tune the model and incur a compute cost, or do not fine-tune the model and discard the data. After the decision, the performance of the current model is measured in terms of an application-specific performance metric such as classification accuracy. Our objective is to learn an optimal policy that determines \emph{when to fine-tune the model} on a single task (e.g., sentiment analysis), under a finite compute budget. We formulate this online decision-making problem as a constrained Markov Decision Process, where the system state captures three essential aspects: (\textit{i}) model's performance, (\textit{ii}) computational budget, and (\textit{iii}) data distribution relevance to historic data encountered up to that point. The transition to the next state is stochastic and therefore, we propose a reinforcement learning-based method to solve this problem, namely the \emph{actor-critic} algorithm. We also consider the special case where the performance of fine-tuning for a given model can be predicted or estimated prior to decision; in this case the problem becomes a Dynamic Programming one. Experiments with a large pre-trained model on a widely-used text classification dataset demonstrate that our method consistently outperforms fine-tuning approaches with the same compute budget by more than $4\%$ in terms of accuracy and achieves $97\%$ of full-parameter fine-tuning accuracy while requiring only $25\%$ of the fine-tuning steps.
Comments6 pages, 2 figures, 1 table, accepted and presented at ICMLCN 2026