arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15990cs.CLcs.LG

少样本退化并非表面所见:跨12个模型、2个任务和2种架构的行为证据、表示分析与随机文本对照

Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 Architectures

Volodymyr Ovcharov

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过行为实验和表示分析发现,少样本提示的性能退化主要源于提示长度而非示例内容,并提出内容增量指标以准确预测模型获益,因果实验证实了该结论。

中文摘要 AI 辅助

少样本提示有时会降低语言模型的性能而非帮助它们,但原因尚不清楚。我们在两个乌克兰语任务——新闻分类和法律案件结果预测——上评估了12个开放权重模型,发现该效应强烈依赖于任务:同一批模型在新闻任务上平均提升24个百分点,在法律文本上仅提升3.4个百分点,且有两个模型出现性能退化。为理解原因,我们深入模型内部。先前工作衡量零样本与少样本模式之间隐藏状态的变化量,但少样本提示要长得多,而仅长度差异就会移动表示。我们提出一个简单修正:用长度匹配的随机文本替换示例,以衡量提示长度引起的位移,然后将其减去。所得指标“内容增量”隔离出模型表示因示例内容而非长度而发生的变化。这彻底改变了图景:原始位移无法预测少样本是有益还是有害(r = 0.20),但内容增量可以(rho = +0.65,p = 0.043)。那些因示例内容而更大幅度重构表示的模型获益更多——这与直觉上的“扭曲”解释相反。在Llama 3.3 70B中掩蔽示例因果性地证实了这一发现,恢复的准确率高于零样本基线。

英文摘要

Few-shot prompting sometimes degrades language models instead of helping them, but why this happens is unknown. We evaluate 12 open-weight models on two Ukrainian tasks news classification and legal case outcome prediction and find that the effect is strongly task-dependent: the same models that gain +24 pp on news show only +3.4 pp on legal text, with two models degrading. To understand why, we look inside the models. Prior work measures how much hidden states shift between zero-shot and few-shot modes, but few-shot prompts are much longer, and that length difference alone moves representations. We propose a simple fix: replace demonstrations with length-matched random text to measure the shift caused by prompt length, then subtract it. The resulting metric content delta isolates how much the model's representations change because of what the demonstrations say, not how long they are. This changes the picture entirely: raw shift does not predict whether few-shot helps or hurts (r = 0.20), but content delta does (rho = +0.65, p = 0.043). Models that restructure representations more from demonstration content benefit more the opposite of the intuitive "distortion" explanation. Masking demonstrations in Llama 3.3 70B confirms the finding causally, recovering accuracy above the zero-shot baseline.

发表机构

  • LEX AI Platform, legal.org.ua(LEX AI平台,legal.org.ua)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑