arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20371cs.CLcs.AI

当大语言模型(LLM)取代微调的自然语言理解(NLU)模型?面向生产型对话系统意图检测的决策框架

When Do LLMs Replace Fine-Tuned NLU? A Decision Framework for Intent Detection in Production Conversational Systems

Carson Rodrigues, Oysturn Vas

首次发表
浏览论文内容

中文总结 AI 辅助

该研究验证LLM能否取代微调NLU模型用于意图检测,对比多方法后得出结论并提炼出生产型对话系统意图检测的决策框架。

中文摘要 AI 辅助

常见观点认为,零样本大语言模型(LLM)可取代用于意图检测的微调NLU分类器。我们直接验证这一观点,发现答案取决于意图空间。在完整ATIS和CLINC150数据集上,我们对比了微调RoBERTa、TF-IDF+逻辑回归基线、句子嵌入k近邻(kNN)以及Claude Haiku零样本方法,报告了自助法95%置信区间和配对显著性检验结果。当存在大量领域内标签时,微调RoBERTa的性能相当或更优,且成本和速度低三个数量级:在ATIS数据集上,其比Claude零样本方法高出11.8个百分点(95.9 vs. 84.1,p<0.001);在包含150个意图的CLINC150广义意图架构上,两者性能统计上无显著差异(89.1 vs. 88.5,p=0.24),即LLM无需训练数据即可匹配全监督模型的性能。LLM的优势体现在三种与生产相关的场景:超出范围检测(OOS召回率为85.6,而RoBERTa为58.1);通过受控的文本-语音转换加噪声再加Whisper pipeline实现对现实自动语音识别(ASR)噪声的鲁棒性(0 dB时为92.5,而RoBERTa为80.0);以及动态部署架构,即针对某一应用意图训练的分类器在新应用意图上得分为0%,而基于架构提示的LLM无需重新训练即可在两个应用上达到约94%的性能。我们将这些发现提炼为面向从业者的决策框架。

英文摘要

A common claim is that zero-shot large language models (LLMs) can replace fine-tuned NLU classifiers for intent detection. We test this claim head-to-head and find that the honest answer is: it depends on the intent space. On full ATIS and CLINC150 we compare a fine-tuned RoBERTa, a TF-IDF+logistic-regression baseline, sentence-embedding kNN, and Claude Haiku zero-shot, reporting bootstrap 95% confidence intervals and paired significance tests. When abundant in-domain labels exist, fine-tuned RoBERTa is as good or better and three orders of magnitude cheaper and faster: on ATIS it beats Claude zero-shot by 11.8 points (95.9 vs. 84.1, p<0.001). On the broad 150-intent CLINC150 schema the two are statistically tied (89.1 vs. 88.5, p=0.24): the LLM matches a fully supervised model with no training data. The LLM's advantages appear in three production-relevant regimes: out-of-scope detection (OOS recall 85.6 vs. 58.1 for RoBERTa); robustness to realistic ASR noise via a controlled text-to-speech to noise to Whisper pipeline (92.5 vs. 80.0 at 0 dB); and dynamic per-deployment schemas, where a classifier trained on one app's intents scores 0% on a new app's intents while the schema-prompted LLM serves both at ~94% with zero retraining. We distill these findings into a decision framework for practitioners.

发表机构

  • Celabe
  • University of Waterloo(滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑