基准测试系统一决策模型与训练分类器及语言模型在自动化决策门控中的表现
Benchmarking System One decision models against trained classifiers and language models for automated decision gates
浏览论文内容
中文总结 AI 辅助
本研究基准测试系统一决策模型、训练分类器及语言模型在自动化决策门控中的表现,发现模型排名依赖条件,并提出了条件依赖的设计规则。
中文摘要 AI 辅助
将分支决策交给模型的软件需要一个声明的选项及其可设定阈值的概率。类型化决策模型,也称为系统一模型,无需生成文本即可返回此类概率,而监督分类器和生成式语言模型则是既有的替代方案。在匹配条件下,一个测试框架向来自六个家族的八个决策模型检查点(包括托管模型Jev)以及两个生成式比较器发送相同的语义请求,并在相同的工作流、意图和社会科学条目上对监督分类器和零样本分类器进行评分。模型类别的排名取决于条件。使用任务自身的标签时,小型训练分类器在意图上最为准确,在工作流上与最佳决策模型无显著差异。无标签时,除基于编码器的检查点外,每个决策模型在工作流和意图上均超过零样本蕴含分类器。通过选项键似然解读,较大的生成模型在工作流和意图上与Jev持平,并在5%风险下接受更多工作流决策,而微调的决策检查点在意图准确性上优于其未调优的基础模型。在少量选项上拟合的存储温度在大量选项上会提高校准误差,而针对5%范围内风险的保留阈值仍使Jev接受了0.310的范围内请求。交换“是”和“否”会使Jev每百个答案翻转50.5个,而微调的检查点减少了其基础模型在社会科学条目上的翻转。一个经过意图训练的第一阶段升级到Jev,在全图形处理器利用率下,以0.43的成本达到其准确性。这些结果产生了条件依赖的自动化决策门控设计规则。
英文摘要
Software that hands branching decisions to a model needs a declared option and a probability it can threshold. Typed decision models, also called System One models, return such probabilities without generating text, while supervised classifiers and generative language models are the established alternatives. One harness sends eight decision-model checkpoints from six families, including the hosted model Jev, and four open generative models from three developers the same semantic requests, and scores trained and zero-shot classifiers on the same workflow, intent, emotion and social-science items. With task labels, a fine-tuned DeBERTa-v3-large has the highest observed accuracy on every labeled benchmark but one. Without labels, no decision model is significantly more accurate than Jev on workflows or intents, but Gemma-4-31B matches it on workflows and exceeds it on CLINC-150 at higher cost and latency. Stated probabilities of generative models become unreadable when replies miss the key format, whereas key likelihoods avoid this but can saturate. A guaranteed 5 percent risk leaves Jev 0.528 of the intent decisions, and an in-scope threshold still accepts 0.310 of out-of-scope requests. On typed-decisions, swapping yes and no flips 50.5 answers per hundred for Jev and at least 16.8 for every generative model tested, against at most 6.5 for four fine-tuned decision checkpoints. Exposure to a benchmark's training data explains the largest lead of an open checkpoint, which vanishes on rater-labeled emotions. An intent-trained first stage escalating to Gemma-4-31B reaches that model's accuracy at about Jev's price. The results yield condition-dependent design rules for automated decision gates.
发表机构
- Texas State University(德克萨斯州立大学)
机构由 AI 辅助整理,请以论文原文为准。