arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13866cs.LGcs.CL

用于小样本文本分类的大语言模型生成样本的几何过滤

Geometric Filtering of LLM-Generated Samples for Few-Shot Text Classification

Benjamín Schindler, Gonzalo A. Ruz

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对LLM生成样本质量异质性问题,提出基于句子嵌入空间欧氏距离的几何过滤框架,结合软加权机制提升小样本文本分类性能,在多任务、多模型及多配置下均表现优异。

中文摘要 AI 辅助

大语言模型(LLM)可生成用于文本分类的合成训练数据,但生成样本的质量存在异质性:部分样本处于嵌入空间的正确类别区域,另一些则落在边缘或跨类别区域。我们提出一种几何过滤框架,在句子嵌入空间中通过每个LLM生成样本与真实类别示例的欧氏距离来评估样本,仅选择几何上一致的候选样本。该框架采用软加权机制将过滤得分转换为分类器训练的样本权重。在13个数据集、5个分类器、10种数据增强方法及超6700种配置上进行评估,我们的方法较SMOTE提升了2.61个百分点(pp),具有统计学显著性(p<0.0001,Cohen's d=0.95,胜率为88.9%)。该方法无需修改过滤器即可推广到命名实体识别任务,性能提升9.26pp,胜率达100%,且对来自4家供应商的5种LLM均表现稳健。一项关键发现是,最简单的基于距离的过滤器始终优于复杂的多准则替代方案。

英文摘要

Large language models (LLMs) can generate synthetic training data for text classification, but the quality of generated samples is heterogeneous: some fall in correct class regions of the embedding space while others land in peripheral or cross-class zones. We propose a geometric filtering framework that evaluates each LLM-generated sample by its Euclidean distance to real class examples in a sentence embedding space, selecting only geometrically consistent candidates. A soft weighting mechanism transforms filter scores into sample weights for classifier training. Evaluated across 13 datasets, 5 classifiers, 10 augmentation methods, and over 6,700 configurations, our method achieves +2.61 percentage points (pp) over SMOTE ($p<0.0001$, Cohen's $d=0.95$, 88.9% win rate). The approach generalizes to named entity recognition (+9.26pp, 100% win rate) without filter modification, and is robust across 5 LLMs from 4 providers. A key finding is that the simplest distance-based filter consistently outperforms complex multi-criteria alternatives.

补充信息

↑