arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLMs 还是朴素贝叶斯?旧宝石还是新方法

LLMs or Naive Bayes? Old Gems or New Ways

Mohammad Firas Sada, Dmitry Mishin, John Graham, Seungmin Kim, Mahidhar Tatineni, Frank Würthwein

arXiv 2609.13185首次发表:更新:

发表机构

San Diego Supercomputer Center, University of California, San Diego; Lawrence Berkeley National Laboratory; Yonsei University College of Medicine(加州大学圣迭戈分校圣迭戈超级计算机中心; 劳伦斯伯克利国家实验室; 延世大学医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对比朴素贝叶斯与多规模LLMs在文本分类中的表现,发现NB在有标注数据时性价比更高,并提供自动模型选择的Kubernetes工具。

AI 中文摘要

大型语言模型(LLMs)在研究计算中引发了一个反复出现的问题:像朴素贝叶斯(NB)这样的经典方法是否应该被淘汰?我们将 Complement Naive Bayes 与跨四个模型家族、规模跨度达 37 倍(从 27B 到 1T 参数的混合专家模型)的零样本和少样本 LLMs 在文本分类任务上进行了基准测试。LLMs 仅在零数据场景中占主导地位(在 Amazon Polarity 情感分析上为 98.0% 对比 88.2%),而且即使这种优势也容易受到污染影响:在低污染情感任务上,NB 击败了零样本 LLM(81.7% 对比 73.0%)。然而,一旦有标注数据可用(例如 AG News),NB 达到 89.1% 的准确率,与零样本 27B LLM(89.0%)在统计上无显著差异,并且优于 397B 前沿模型(84.8%),同时在普通 CPU 上以每秒数千个样本的速度运行。微调后的 DistilBERT 达到 90.6%,但在批大小为 1 时吞吐量远低于 NB(表 2)。我们测量的 GPU 吞吐量分析显示,小 LLM 的批处理推理比 NB 的 CPU 推理慢 40-486 倍(倍数强烈依赖于主机 CPU),这暴露了一个受内存带宽限制的结构性差距,每个样本的能耗大约低两个数量级。对于资源受限的 HPC 从业者,在带有标注数据的文本分类任务中,NB 仍然是最优选择。我们表明决策线是任务相关的(NB 在主题分类中大约在 N ~ 10^4 个标签时达到 LLM 水平,而零数据情感分析在所有测试的 N 下都偏向 LLM),并提供了一个 Kubernetes Helm 操作符,使用可配置阈值和可验证的 Prometheus 指标来自动化模型选择。

英文摘要

Large language models (LLMs) prompt a recurring question in research computing: should classical methods like Naive Bayes (NB) be retired? We benchmark Complement Naive Bayes against zero-shot and few-shot LLMs spanning four model families and a 37x range in scale (27B to a 1T-parameter mixture-of-experts) across text classification tasks. LLMs dominate only in zero-data regimes (98.0% vs 88.2% on Amazon Polarity sentiment), and even that win is contamination-prone: on a low-contamination sentiment task NB beats the zero-shot LLM (81.7% vs 73.0%). However, once labeled data is available (e.g., AG News), NB reaches 89.1% accuracy, statistically indistinguishable from the zero-shot 27B LLM (89.0%) and better than the 397B frontier model (84.8%), at thousands of samples/sec on a commodity CPU. Fine-tuned DistilBERT reaches 90.6% but at far lower throughput than NB at batch size 1 (Table 2). Our measured GPU throughput analysis shows small-LLM batched inference is 40-486x slower than NB CPU inference (the multiplier depends strongly on the host CPU), exposing a structural gap bounded by memory bandwidth, with roughly two orders of magnitude lower energy per sample. For resource-constrained HPC practitioners performing text classification with labeled data, NB remains the optimal choice. We show the decision line is task-dependent (NB reaches LLM parity around $N \sim 10^4$ labels for topic classification, while zero-data sentiment favors the LLM at all N tested) and provide a Kubernetes Helm operator that automates model selection using configurable thresholds and verifiable Prometheus metrics.

Comments6 pages, 3 figures, 5 tables. Accepted to the Proceedings of Practice and Experience in Advanced Research Computing (PEARC '26), Minneapolis, Minnesota, USA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑