arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22969cs.IR

少样本提示何时起作用?跨模型规模、架构和输出解析鲁棒性的样本数量效应的系统实证研究

When Does Few-Shot Prompting Help? A Systematic Empirical Study of Shot-Count Effects Across Model Scale, Architecture, and Output Parsing Robustness

Ayush Dwivedi, Ashvi Soni

中文总结 AI 辅助

研究少样本提示中样本数量与模型规模、架构及输出格式合规性对分类性能的影响。通过对五个LLM在六种样本数量配置下研究,发现样本数量与分类性能关系非单调、非普遍且不能仅由模型规模预测,还纠正了影响Llama 3.3 70B性能的解析工件。

中文摘要 AI 辅助

少样本提示是在向大语言模型(LLM)呈现查询之前,在其前面添加少量输入-输出示例对的做法,是自然语言处理中最广泛采用的推理时技术之一。然而,很少有系统的工作研究样本数量在确定分类性能时如何与模型规模、架构和输出格式合规性相互作用。本文对五个LLM在AG新闻四类基准(n = 200)上的六种样本数量配置(k ∈ {0,1,2,3,5,8})进行了对照研究。我们的模型涵盖专有和开源系列:Gemini Flash Lite、GPT-4o-mini、Llama 3.1 8B、Llama 3.3 70B和Llama 4 Scout 17B。我们报告了所有30种配置下的宏观平均F1以及95%的自举置信区间(B = 10,000)、置换检验p值和科恩d效应大小。我们的发现揭示了四种质的不同行为模式:(1)在零样本时已经校准良好的模型,显示出适度的、统计上不显著的增益(Gemini、GPT-4o-mini);(2)经历灾难性零样本失败但用单个示例显著恢复的模型(Llama 3.1 8B,d = 10.98,p < 0.0001);(3)在零样本时最优但随着额外示例单调退化的模型(Llama 4 Scout);(4)呈现U形曲线的模型(Llama 3.3 70B:零样本F1 = 0.907,2样本F1 = 0.635,5样本F1 = 0.785,解析器校正后)。我们还识别、诊断并纠正了一个系统性解析工件,该工件人为地使Llama 3.3 70B的性能降低了高达206%,这是对LLM评估实践的方法学贡献。我们的结果表明,样本数量与分类性能之间的关系不是单调的、不是普遍的,也不能仅从模型规模预测。

英文摘要

Few-shot prompting, the practice of prepending a small number of input-output demonstration pairs to a query before presenting it to a large language model (LLM), is among the most widely adopted inference-time techniques in NLP. Yet little systematic work investigates how shot count interacts with model scale, architecture, and output format compliance in determining classification performance. This paper presents a controlled study of five LLMs across six shot-count configurations (k in {0,1,2,3,5,8}) on the AG News four-class benchmark (n=200). Our models span proprietary and open-source families: Gemini Flash Lite, GPT-4o-mini, Llama 3.1 8B, Llama 3.3 70B, and Llama 4 Scout 17B. We report macro-averaged F1 with 95% bootstrap confidence intervals (B=10,000), permutation-test p-values, and Cohen's d effect sizes across all 30 configurations. Our findings reveal four qualitatively distinct behavioral regimes: (1) models already well-calibrated at zero-shot that show modest, statistically insignificant gains (Gemini, GPT-4o-mini); (2) models that undergo catastrophic zero-shot failure but recover dramatically with a single example (Llama 3.1 8B, d=10.98, p<0.0001); (3) models optimal at zero-shot that degrade monotonically with additional examples (Llama 4 Scout); and (4) models exhibiting a U-shaped curve (Llama 3.3 70B: 0-shot F1=0.907, 2-shot F1=0.635, 5-shot F1=0.785 with parser corrected). We additionally identify, diagnose, and correct a systematic parsing artifact that artificially deflated Llama 3.3 70B performance by up to 206%, constituting a methodological contribution to LLM evaluation practice. Our results demonstrate that the relationship between shot count and classification performance is not monotonic, not universal, and not predictable from model scale alone.

补充信息

↑