arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

监督微调何时会降低指令敏感性?

When Does Supervised Fine-Tuning Reduce Instruction Sensitivity?

Jaekeol Choi

arXiv 2608.26661首次发表:更新:

AI 中文总结

该研究探讨了监督微调(SFT)对大型语言模型指令敏感性的影响,发现SFT的稳健性效应因模型规模、类型及适应设置而异,且敏感性还受预测评分协议影响。

AI 中文摘要

大型语言模型在同一任务指令的不同表述下会表现出显著的性能差异,然而传统的特定任务监督微调(SFT)如何改变这种指令敏感性仍不清楚。我们通过在多个改写后的指令下评估固定模型检查点,并将指令敏感性定义为任务性能在这些指令间的标准差,来研究该问题。我们在MS MARCO数据集上对1.7B、4B和8B规模的Qwen3模型开展了受控规模分析,同时使用Mistral-7B和Gemma-2-9B进行了针对性的跨系列验证。在SFT之前,Qwen3模型的指令敏感性随规模增大急剧下降;在1.7B和4B规模下,SFT一致降低了训练指令间的敏感性,降幅约为54%至71%;在8B规模下,单个敏感性变化与零无统计学差异,但在查询级自助法分析下,训练指令间的配对对比具有统计学可靠性,且在所有三个随机种子上方向一致。Gemma-2-9B展现出与Qwen3-8B相同的训练指令对比方向,而Mistral-7B则没有,表明该效应的强度也因模型而异。在ESCI-English上的实验进一步显示,即使有效标签生成近乎完美且平均任务性能相似,自由生成和基于似然的强制选择评估也可能得出性质不同的稳健性结论。总体而言,SFT并未统一降低指令敏感性:其稳健性效应取决于适应设置,而测得的敏感性还可能取决于预测和评分协议。

英文摘要

Large language models can exhibit substantial performance variation across alternative formulations of the same task instruction, yet it remains unclear how conventional task-specific supervised fine-tuning (SFT) changes this instruction sensitivity. We study this question by evaluating fixed model checkpoints under multiple paraphrased instructions and defining instruction sensitivity as the standard deviation of task performance across them. We conduct a controlled scale analysis with Qwen3 models at 1.7B, 4B, and 8B on MS MARCO, together with targeted cross-family checks using Mistral-7B and Gemma-2-9B. Before SFT, instruction sensitivity decreases sharply with Qwen3 model scale. At 1.7B and 4B, SFT consistently reduces sensitivity across training instructions, with reductions of approximately 54--71%. At 8B, individual sensitivity changes are not statistically distinguishable from zero, but paired contrasts between training instructions are statistically reliable under query-level bootstrap analysis and have consistent directions across all three random seeds. Gemma-2-9B shows the same directional training-instruction contrast as Qwen3-8B, whereas Mistral-7B does not, suggesting that the strength of this effect also varies across models. Experiments on ESCI-English further show that free-generation and likelihood-based forced-choice evaluation can yield qualitatively different robustness conclusions even when valid-label generation is nearly perfect and average task performance is similar. Overall, SFT does not uniformly reduce instruction sensitivity: its robustness effect depends on the adaptation setting, while measured sensitivity can additionally depend on the prediction and scoring protocol.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑