发表机构
ufak AI(ufak AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HakemBench是土耳其语类型化决策基准,含2,346个条目和七个轨道,评估模型决策质量、校准与选择性自动化,并报告探测结果;排行榜领先者综合得分0.888。
AI 中文摘要
HakemBench是一个土耳其语类型化决策基准,其中被测模型阅读一段文本、一个问题以及一组固定的选项,并为每个选项返回一个概率。1.0版本在CC BY 4.0许可下完全开放发布,包含2,346个条目和4,275个选择、是非及评分问题,涵盖七个轨道(事实核查分流、教育、护栏、法律路由、内容审核、垃圾邮件与钓鱼、客户支持)。一个评估框架对决策质量(宏F1)、校准(基于归一化Brier分数)和选择性自动化(基于广义风险覆盖曲线的归一化面积)进行评分,通过几何平均合并这些指标,并报告来自2,000次自助抽样(bootstrap draws)的区间;同时报告针对选项顺序、改写、英文翻译和替换名称的探测结果。大多数黄金标签来自一个AI模型系列的盲测结果,与其他模型系列的大型语言模型评审团投票进行比较;这些标签未经人工验证。在16行的排行榜上,领先者综合得分为0.888,而实验室自己的模型以0.660排名第7。其数字并非盲测。早期运行的测试结果影响了其训练数据,因此其护栏、内容审核和客户支持数字被标记;当每个模型仅在其他四个轨道上评分时,其综合得分为0.678,在16个中排名第6。
英文摘要
HakemBench is a Turkish benchmark of typed decisions, in which the model under test reads a text, a question and a fixed set of options and returns a probability for every option. Version 1.0 is released fully open under CC BY 4.0, with 2,346 items and 4,275 choice, yes/no and score questions in seven tracks (fact-check triage, education, guardrails, legal routing, moderation, spam and phishing, and customer support). One harness scores decision quality (macro F1), calibration (from the normalised Brier score) and selective automation (from the normalised area under the generalised risk-coverage curve), combines them by a geometric mean and reports intervals from 2,000 bootstrap draws; probes for option order, paraphrase, English translation and substituted names are reported alongside. Most gold labels come from blind passes of one AI model family compared with the votes of a panel of large language models from other model families; they are not human-verified. On a board of 16 rows the leader scores a composite of 0.888 and the lab's own model is 7th at 0.660. Its numbers are not blind. Earlier runs' test results shaped its training data, so its guardrail, moderation and customer support numbers are flagged; with every model scored on the other four tracks only, its composite is 0.678, 6th of 16.
Comments9 pages including references. Data, harness, scorer and board: https://huggingface.co/datasets/ufakai/HakemBench, https://github.com/ufakai/hakembench