arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

克服先验障碍:长尾分布下的监督微调

Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution

Haohui Wang, Jiahao Xu, Wangzhi Zhan, Tong Zeng, Dongqi Fu, Hong Li, Swastik Roy, Naren Ramakrishnan, Chris North, Jian Kang, Yujun Yan, Dawei Zhou

arXiv 2610.12345首次发表:更新:

发表机构

Virginia Tech; Amazon; Meta; MBZUAI; Dartmouth College(弗吉尼亚理工大学; 亚马逊公司; Meta; 穆罕默德·本·扎耶德人工智能大学; 达特茅斯学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对长尾分布下SFT中尾部概念先验障碍较高的问题,提出自适应指令选择方法PASS,在有限预算下提升了SFT性能,优于七种现有方法。

AI 中文摘要

监督微调(SFT)可将预训练大语言模型(LLM)适配至下游任务,但所需概念获得的预训练支持程度存在显著差异:常见概念更易被充分学习,而稀有概念可能仍表征不足。我们引入名为“先验障碍”的新概念,用于量化预训练模型对竞争概念的支持强度相对于目标概念的程度。我们观察到,先验障碍遵循长尾分布,这使得头部概念与尾部概念在SFT中处于不同的起点:头部概念面临较低的先验障碍,而尾部概念则需要额外指令来克服其较高的先验障碍。我们的理论分析进一步推导了长尾先验障碍下SFT的预测风险界,明确刻画了先验障碍与累积SFT证据如何共同决定预测性能。受此依赖先验障碍的需求驱动,我们提出PASS,一种自适应SFT指令选择方法,该方法构建源自参考的概念,估计每条指令提供的区分证据,并在当前选择下向覆盖不足的概念自适应分配选择预算。通过这种方式,PASS在有限预算下同时考虑哪些指令可提供有用证据,以及何处需要额外监督。实验表明,在四种主干-预算设置下,我们的方法始终优于七种最先进的指令选择方法;消融研究进一步显示,PASS的自适应分配始终优于均匀分配。

英文摘要

Supervised fine-tuning (SFT) adapts pretrained large language models (LLMs) to downstream tasks, but the required concepts can receive substantially different levels of pretrained support. Frequent concepts are more likely to be well learned, whereas rare concepts may remain weakly represented. We introduce a novel notion named prior barrier to quantify how strongly the pretrained model supports competing concepts over the target concept. We observe that prior barriers follow a long-tail distribution, placing head and tail concepts at different starting points for SFT: head concepts face lower prior barriers, whereas tail concepts require additional instructions to overcome their higher prior barriers. Our theoretical analysis further derives a predictive risk bound for SFT under long-tail prior barriers, explicitly characterizing how the prior barrier and accumulated SFT evidence jointly determine predictive performance. Motivated by this prior barrier-dependent demand, we propose PASS, an adaptive SFT instruction selection method that constructs reference-derived concepts and estimates the distinguishing evidence provided by each instruction, and adaptively allocates the selection budget toward concepts that remain insufficiently covered under the current selection. In this way, PASS jointly considers which instructions can provide useful evidence and where additional supervision is needed under a limited budget. Experiments show that our method consistently outperforms seven state-of-the-art instruction selection methods on four backbone-budget settings. An ablation study further shows that PASS's adaptive allocation consistently improves over uniform allocation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑