ufakzeka-karar:一种具有顺序不变选项评分的开放土耳其语类型化决策模型
ufakzeka-karar: An Open Turkish Typed-Decision Model with Order-Invariant Option Scoring
浏览论文内容
中文总结 AI 辅助
本文提出 ufakzeka-karar,一个 1.82 亿参数的开放土耳其语决策模型,通过顺序不变的选项评分和温度缩放概率实现高效推理,在 HakemBench 上取得有竞争力的性能,并公开权重与代码。
中文摘要 AI 辅助
ufakzeka-karar 是一个开放的土耳其语决策模型,拥有 182,494,466 个参数。给定一段土耳其语文本和固定答案类型的问题(选择、有序量表上的等级或是否),该模型为每个选项返回温度缩放后的概率,以及一个作为“不确定”信号的预期误差,无需生成文本,并且对于最多十个选项只需一次 CPU 前向传播。该模型基于实验室的 ufakzeka-1-base,其头部在共享位置上独立于其他选项对每个选项进行评分,因此答案不依赖于选项顺序。使用打乱选项训练的序列头部在准确率上与之相当,但当仅选项顺序改变时,其答案中有 2.3% 到 2.8% 发生变化;REINFORCE 在宏 F1 上比交叉熵损失低 10.2 个百分点(0.102)。在 HakemBench v1.0 的开放集(4,275 个问题,7 个轨道)上,发布的模型在 16 行中排名第 7,综合得分为 0.660(95% 置信区间 0.642 至 0.677)。温度缩放降低了开发集上的校准误差(平滑 ECE),但在未见过的支持问题上反而提高了校准误差,对于第一个评分运行,从 0.027 升至 0.045,而该运行从未在这些问题上训练过;发布的模型后来在这些问题上进行了训练,因此其 0.036 至 0.064 的结果并非未见问题的测试。发布的模型是在 HakemBench 上评分的三次运行中的最后一次,其数字并非盲测。第二次运行的新训练数据针对第一次运行在完整测试集的护栏、审核和客户支持轨道上的错误,而发布的运行是在阅读了第二次运行在完整测试集上的护栏结果之后训练的,其协议在其任何数据、代码或运行之前以书面形式固定。其所有数字均在这些阅读之后得出;其护栏、审核和客户支持数字带有“受阅读测试结果影响”的标记。当每个模型仅在其他四个轨道上评分时,其综合得分为 0.678,在 16 个中排名第 6。权重和代码采用 Apache-2.0 许可证。
英文摘要
ufakzeka-karar is an open Turkish decision model with 182,494,466 parameters. Given a Turkish text and questions of a fixed answer type (a choice, a level on an ordered scale, or yes or no), it returns a temperature-scaled probability for every option and an expected error that serves as a "not sure" signal, without generating text and in one CPU forward pass for up to ten options. Built on the lab's ufakzeka-1-base, its head scores each option blind to the others at shared positions, so the answer does not depend on option order. A sequential head trained with shuffled options was about as accurate but changed 2.3 to 2.8 percent of its answers when only the option order changed; REINFORCE lost 10.2 points (0.102) of macro F1 to cross-entropy. On the open set of HakemBench v1.0 (4,275 questions, 7 tracks) the released model ranks 7th of 16 rows with a composite of 0.660 (95% interval 0.642 to 0.677). Temperature scaling lowers calibration error (smooth ECE) on the development set but raises it on held-out support questions, from 0.027 to 0.045 for the first scored run, which never trained on them; the released model later trained on them, so its 0.036 to 0.064 is not an unseen-question test. The released model is the last of three runs scored on HakemBench, and its numbers are not blind. The second run's new training data was aimed at the first run's errors on the full test set in guardrails, moderation and customer support, and the released run was trained after the second run's guardrail results on the full test set were read, under a protocol fixed in writing before any of its data, code or runs. All its numbers come after these readings; its guardrail, moderation and customer support numbers carry the flag "shaped by reading the test results". With every model scored on the other four tracks only, its composite is 0.678, 6th of 16. Weights and code are under Apache-2.0.
发表机构
- ufak AI
机构由 AI 辅助整理,请以论文原文为准。