平衡k-shot采样下的精确简并:对LLM嵌入小样本判别分析的影响
Exact Degeneracy Under Balanced k-Shot Sampling:Consequences for Small-Sample Discriminant Analysis on LLM Embeddings
查看机构详情
- The University of Aizu(会津大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文证明平衡k-shot采样导致KLPCDA小样本判别分析精确简并,提出公式内平局决胜修复,并在LLM嵌入少样本分类中验证修复效果,发现逻辑回归探针仍更优,差距主要源于估计效率。
中文摘要 AI 辅助
平衡k-shot采样每类恰好抽取k个标记样本。我们证明,它在一族小样本判别估计器中引发一种精确的、可证明的简并。在平衡采样下,核化线性主成分判别分析(KLPCDA)的类内散度算子不仅仅是秩亏的,而是精确地成为一个缩放的正交投影算子。我们以闭式推导其后果:KLPCDA七个变体中的两个,其所有信号特征值完全相等,因此其特征向量选择准则可证明是无关紧要的,而非病态的;第三个变体则具有可证明的空目标。这源于估计器的构造,而非任何数据集;我们在冻结的句子嵌入上,以及分别在一个仅解码器生成模型的残差流激活上,确认了这一点。公式内的平局决胜修复了那两个可修复的变体,其恢复受类别数量门控:在少类别数据集上,残差子空间约束的代价是多类别数据集的5倍(p=0.000001)。随后,我们在冻结的LLM嵌入(n远小于d,最高达4096)上的少样本文本分类中评估修复后的框架,涵盖四个数据集、三种嵌入尺寸和三个训练基线(SetFit、LoRA、上下文学习)。一个经过适当交叉验证的逻辑回归探针在四个数据集中的三个上,在每种嵌入尺寸下,仍击败所有KLPCDA变体;从像素、振动信号和基因表达数据中获得的指导并不能直接推广到这一特征空间。三个独立的几何可分性指标无法解释为何一个高维基于解码器的嵌入模型不如较小的双向编码器,排除了各向异性;这一差距实质上是一种估计效率效应,而非永久上限,当支持集从k≤10增长到k=30-50时,差距缩小超过80%(p=0.00195,两个多类别数据集均如此)。
英文摘要
Balanced k-shot sampling draws exactly k labeled examples per class. We show that it induces an exact, provable degeneracy in a family of small-sample discriminant estimators. Under balanced sampling, the within-class scatter operator of Kernelized Linear Principal Component Discriminant Analysis (KLPCDA) is not merely rank-deficient but exactly a scaled orthogonal projector. We derive the consequences in closed form: two of KLPCDA's seven variants have every signal eigenvalue exactly equal, so their eigenvector selection criterion is provably indifferent rather than ill-conditioned, and a third has a provably void objective. This follows from the estimators' construction, not any dataset; we confirm it on frozen sentence embeddings and, separately, on residual-stream activations from a decoder-only generative model. An in-formula tie-break repairs the two repairable variants, with recovery gated by class count: the residual subspace constraint costs 5x more on few-class than many-class datasets (p=0.000001). We then evaluate the repaired framework on few-shot text classification on frozen LLM embeddings (n much smaller than d, up to 4096), across four datasets, three embedding sizes, and three trained baselines (SetFit, LoRA, in-context learning). A properly cross-validated logistic-regression probe still beats every KLPCDA variant on three of four datasets, at every embedding size; guidance carried from pixel, vibration-signal, and gene-expression data does not directly generalize to this feature space. Three independent geometric separability metrics fail to explain why one high-dimensional decoder-based embedding model underperforms smaller bidirectional encoders, ruling out anisotropy; the gap is substantially an estimation-efficiency effect, not a permanent ceiling, closing by more than 80% when the support set grows from k<=10 to k=30-50 (p=0.00195, both many-class datasets).