发表机构
University of Bonn; Lamarr-Institute for Machine Learning and Artificial Intelligence; Fraunhofer IAIS(波恩大学; 拉马尔机器学习和人工智能研究所; 弗劳恩霍夫智能分析和信息系统研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文评估并基准测试了商业 System One 模型 Jev,发现其在多个任务上表现优异,优于开放模型,且概率校准良好,但低资源语言和噪声标签下性能下降。
AI 中文摘要
Jev 是 TypeSafe AI 推出的商业 System One 模型,它不生成文本:给定一个状态和类型化问题,它从固定选项中选择一个答案、在评分标准上给出一个位置,或返回一个陈述为真的概率,供应商称这些概率是经过校准的。此类模型针对信息访问管道中的小决策,例如查询路由、接地检查、内容审核或按评分标准评分。我们在 37 个数据集上对 Jev(jev-1.13.0)进行零样本评估,涵盖分类、路由、自然语言推理、阅读理解、常识推理、审核、法律条款分析和评分标准评分,每个数据集使用一个冻结模板,并使用完整的评估分割:共 346,009 个请求,花费不到 10 美元。作为参考,我们通过选项上的精确下一个词元概率对 Qwen3.8-27B 和 Gemma-4-E4B 在相同请求上进行评分。Jev 在 IMDB、SST-2、HellaSwag 和 ARC 上达到 95-99% 的准确率,在 Belebele 上跨 122 种语言达到 86.7%。它在 37 个数据集中有 27 个击败 Qwen,且 Qwen 的九个领先优势均未超出自举区间,并在全部 37 个数据集上击败 Gemma。所有三个模型在低资源语言、细粒度或噪声标签以及基于评分标准的质量判断上均表现下降。Jev 的选择概率校准良好,支持选择性预测。二元概率排名良好,但相对于固定的 0.5 阈值定位不佳;在训练数据上调整的阈值将 UNFAIR-ToS 上的微 F1 从 0.50 提高到 0.75。Jev 对 MMLU 中计算密集的问题回答比其他 MMLU 问题更准确(94% 对 91%),而两个开放模型以及所有三个模型在 C-Eval 上都发现这些问题更难。旋转选项不影响 Jev 的准确率,而省略问题则使其准确率降至接近随机水平,这排除了浅层记忆,但不排除记忆的问题-答案对。我们发布了代码、测试框架和所有原始响应。
英文摘要
Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed options, a position on a rubric, or the probability that a statement is true, with probabilities the vendor describes as calibrated. Such models target small decisions in information access pipelines, such as routing queries, checking grounding, moderating content, or rating against a rubric. We evaluate Jev (jev-1.13.0) zero-shot on 37 datasets spanning classification, routing, natural language inference, reading comprehension, commonsense reasoning, moderation, legal clause analysis and rubric scoring, with one frozen template per dataset and full evaluation splits: 346,009 requests for under USD 10. For reference, we score Qwen3.8-27B and Gemma-4-E4B on identical requests via their exact next-token probabilities over the options. Jev reaches 95-99% accuracy on IMDB, SST-2, HellaSwag and ARC and 86.7% on Belebele across 122 languages. It beats Qwen on 27 of 37 datasets, with none of Qwen's nine leads outside the bootstrap intervals, and Gemma on all 37. All three models degrade on low-resource languages, fine-grained or noisy labels, and rubric-based quality judgments. Jev's choice probabilities are well calibrated and support selective prediction. Binary probabilities rank well but are poorly placed relative to a fixed 0.5 threshold; thresholds tuned on training data raise micro-F1 on UNFAIR-ToS from 0.50 to 0.75. Jev answers MMLU's calculation-heavy questions more accurately than other MMLU questions (94% vs. 91%), whereas both open models, and all three on C-Eval, find them harder. Rotating the options leaves Jev's accuracy unchanged and withholding the question drops it to near chance, ruling out shallow memorization but not memorized question-answer pairs. We release the code, harness and all raw responses.
CommentsCode available at github.com/AppliedMachineLearning-Lab/jev-benchmarking, model responses at doi.org/10.5281/zenodo.23039006