arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

专用决策模型与通用大语言模型(LLM):在知识、推理及多语言任务上对Jev进行基准测试

Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks

Xing Li, Qingcheng Chang, Jinzhong Ning, Changfeng Xu, Shenlong Zhang, Yijia Zhang, Ling Luo, Hongfei Lin

arXiv 2610.11978首次发表:更新:

发表机构

Dalian Maritime University; Dalian University of Technology(大连海事大学; 大连理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在13项多选择基准上对比专用决策模型Jev与19个分三级的通用LLM,发现Jev在知识类任务上可媲美前沿级LLM,但在数学应用题上表现逊于所有LLM。

AI 中文摘要

Jev是一种“系统一”模型,它返回给定选项中的一个选择而非生成文本。本研究探讨这类专用决策模型与通用大语言模型(LLM)的对比情况。我们在涵盖知识、推理及多语言理解的13个多项选择基准上对Jev进行评估,并将其与三个层级共19个LLM(前沿级、代表性级及小型级)展开对比。Jev在知识与常识基准上可与前沿级LLM相媲美,且在MMLU-Redux和ARC-Challenge上取得最优分数;在数学领域之外,它还优于大多数代表性级LLM及所有小型级LLM。不过,它在数学应用题上表现落后:在MathQA上,其得分比前沿级中位数低17.7个百分点,且低于全部19个LLM。这些结果表明,专用决策模型在主要依赖知识的决策上可与通用LLM匹敌,但在需要多步计算的决策上则无法做到。

英文摘要

Jev is a "System One" model that returns a choice among given options instead of generating text. We study how such a specialized decision model compares with general-purpose large language models (LLMs). We evaluate Jev on 13 multiple-choice benchmarks covering knowledge, reasoning, and multilingual understanding, and compare it with 19 LLMs in three tiers: frontier, representative, and small. Jev is competitive with frontier LLMs on knowledge and commonsense benchmarks and obtains the best score on MMLU-Redux and ARC-Challenge. Outside mathematics, it also outperforms most representative LLMs and all small LLMs. However, it falls behind on mathematical word problems: on MathQA, it is 17.7 points below the frontier median and scores lower than all 19 LLMs. These results indicate that a specialized decision model can match general-purpose LLMs on decisions that rely mainly on knowledge, but not on decisions that require multi-step calculation.

Comments10 pages, 2 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑