arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27421cs.CLcs.AIcs.LG

选择用于零样本意图分类的开放权重语言模型:41种模型的系统评估

Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models

Parishruthi Ganesh, Gerry Dozier, Cheryl Seals

AI总结:

该研究通过系统评估41种开放权重语言模型,为意图分类场景下选择符合计算、延迟等约束的模型提供了实用指导。

AI中文摘要:

意图分类是面向任务的对话系统的核心组件,但从业者在计算、延迟和鲁棒性约束下,缺乏选择可部署开放权重语言模型的系统指导。我们对15个系列、参数规模在1.35亿至90亿之间的41种开放权重语言模型,在8个英语单标签意图分类数据集上开展了系统的零样本评估。第9个数据集ATIS使用5个带标签演示,作为辅助的五样本结果报告。评估涵盖标准基准、大规模语音助手语料库和生产衍生的电商数据集。除精确匹配准确率外,我们还分析了置信度校准、对真实输入扰动的鲁棒性、模型排名的统计可靠性、部署效率及基准饱和程度。结果显示,经指令微调的30亿参数模型性能可优于多款被评估的70亿参数基础模型;在MASSIVE数据集上,领先模型间的差异经配对McNemar检验后无统计显著性;SNIPS等广泛使用的基准已出现饱和,无法有效区分当前开放权重模型;指令微调对置信度校准的影响不一致,并非均为有害。这些发现为选择和评估用于意图分类的开放权重语言模型提供了实用指导。

英文摘要:

Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployable open-weight language models under compute, latency, and robustness constraints. We present a systematic zero-shot evaluation of 41 open-weight language models spanning 15 families and the 135M--9B parameter range across eight English single-label intent-classification datasets. A ninth dataset, ATIS, uses five labeled demonstrations and is reported as an auxiliary five-shot result. The evaluation includes standard benchmarks, a large-scale voice-assistant corpus, and production-derived e-commerce datasets. Beyond exact-match accuracy, we analyze confidence calibration, robustness to realistic input perturbations, statistical reliability of model rankings, deployment efficiency, and benchmark saturation. Our results show that instruction-tuned 3B models can outperform several evaluated 7B base models, that differences among leading models on MASSIVE are statistically indistinguishable under pairwise McNemar tests, and that widely used benchmarks such as SNIPS have become saturated and no longer meaningfully discriminate among current open-weight models. Instruction tuning's effect on confidence calibration is inconsistent rather than uniformly harmful. These findings provide practical guidance for selecting and evaluating open-weight language models for intent classification.

↑