少样本提示何时起作用?跨模型规模、架构和输出解析鲁棒性的样本数量效应的系统实证研究
When Does Few-Shot Prompting Help? A Systematic Empirical Study of Shot-Count Effects Across Model Scale, Architecture, and Output Parsing Robustness
AI总结:
研究少样本提示中样本数量与模型规模、架构及输出格式合规性对分类性能的影响。通过对五个LLM在六种样本数量配置下研究,发现样本数量与分类性能关系非单调、非普遍且不能仅由模型规模预测,还纠正了影响Llama 3.3 70B性能的解析工件。
AI中文摘要:
少样本提示是在向大语言模型(LLM)呈现查询之前,在其前面添加少量输入-输出示例对的做法,是自然语言处理中最广泛采用的推理时技术之一。然而,很少有系统的工作研究样本数量在确定分类性能时如何与模型规模、架构和输出格式合规性相互作用。本文对五个LLM在AG新闻四类基准(n = 200)上的六种样本数量配置(k ∈ {0,1,2,3,5,8})进行了对照研究。我们的模型涵盖专有和开源系列:Gemini Flash Lite、GPT-4o-mini、Llama 3.1 8B、Llama 3.3 70B和Llama 4 Scout 17B。我们报告了所有30种配置下的宏观平均F1以及95%的自举置信区间(B = 10,000)、置换检验p值和科恩d效应大小。我们的发现揭示了四种质的不同行为模式:(1)在零样本时已经校准良好的模型,显示出适度的、统计上不显著的增益(Gemini、GPT-4o-mini);(2)经历灾难性零样本失败但用单个示例显著恢复的模型(Llama 3.1 8B,d = 10.98,p < 0.0001);(3)在零样本时最优但随着额外示例单调退化的模型(Llama 4 Scout);(4)呈现U形曲线的模型(Llama 3.3 70B:零样本F1 = 0.907,2样本F1 = 0.635,5样本F1 = 0.785,解析器校正后)。我们还识别、诊断并纠正了一个系统性解析工件,该工件人为地使Llama 3.3 70B的性能降低了高达206%,这是对LLM评估实践的方法学贡献。我们的结果表明,样本数量与分类性能之间的关系不是单调的、不是普遍的,也不能仅从模型规模预测。
英文摘要:
Few-shot prompting, the practice of prepending a small number of input-output demonstration pairs to a query before presenting it to a large language model (LLM), is among the most widely adopted inference-time techniques in NLP. Yet little systematic work investigates how shot count interacts with model scale, architecture, and output format compliance in determining classification performance. This paper presents a controlled study of five LLMs across six shot-count configurations (k in {0,1,2,3,5,8}) on the AG News four-class benchmark (n=200). Our models span proprietary and open-source families: Gemini Flash Lite, GPT-4o-mini, Llama 3.1 8B, Llama 3.3 70B, and Llama 4 Scout 17B. We report macro-averaged F1 with 95% bootstrap confidence intervals (B=10,000), permutation-test p-values, and Cohen's d effect sizes across all 30 configurations. Our findings reveal four qualitatively distinct behavioral regimes: (1) models already well-calibrated at zero-shot that show modest, statistically insignificant gains (Gemini, GPT-4o-mini); (2) models that undergo catastrophic zero-shot failure but recover dramatically with a single example (Llama 3.1 8B, d=10.98, p<0.0001); (3) models optimal at zero-shot that degrade monotonically with additional examples (Llama 4 Scout); and (4) models exhibiting a U-shaped curve (Llama 3.3 70B: 0-shot F1=0.907, 2-shot F1=0.635, 5-shot F1=0.785 with parser corrected). We additionally identify, diagnose, and correct a systematic parsing artifact that artificially deflated Llama 3.3 70B performance by up to 206%, constituting a methodological contribution to LLM evaluation practice. Our results demonstrate that the relationship between shot count and classification performance is not monotonic, not universal, and not predictable from model scale alone.