arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KinyaMed:种子,而非行——以错误单位编写的语料需求未能约束什么

KinyaMed: Seeds, Not Rows -- What a Corpus Requirement Written in the Wrong Unit Fails to Constrain

Marius Bayizere

arXiv 2609.32234首次发表:更新:

AI 中文总结

针对基尼亚卢旺达语分诊分类器,发现以行数计量的语料需求无法约束质量,提出以种子短语计数的门控及检测工具,并报告负面结果。

AI 中文摘要

分诊决定谁先被看到。在为患者语音的基尼亚卢旺达语构建紧急程度分类器时,我们发现我们的规范可以在不产生其旨在确保的任何内容的情况下得到满足。我们报告了这一情况,以及检测到这一情况的工具,而不是分类器。针对卢旺达医疗中心接收的四种语言设计,每个工具均按语言划分:句子在四个分支中编写,而行只在一个分支中生成,因为其他三种语言所需的框架槽位不存在。我们的需求要求一百万个示例;生成在130秒内产生了它们。它未能通过该规范中的九个质量门中的四个,当行归因于其源句子时则为六个。绑定门统计不同的已编写种子短语,而非行:在我们的165个种子上,任何行数都无法通过任何语料库。随之而来的是两个不同的不足:还需2,835个句子才能通过种子计数门,还需19,835个才能达到所述的一百万行,因为另一个门将种子限制为50行。行来自机器,每秒7,700行;种子来自临床医生。行数约束了廉价的数量,却让昂贵的数量不受约束,因此根本无法约束质量。还有两个进一步的负面结果。一个由九条不同句子构建的17,942行评估集不支持任何结论:我们的门统计不同句子,拒绝了其38个单元格并报告无结果。一个在统一表面形式的语料库上训练的模型,在大小写变化下对31.5%的输入改变其预测的紧急程度,在单个错字下为21.0%,这作为测量而非归因报告。一个未附带分词器的模型目录加载时无错误,并在无法读取的输入上回答其类别先验,且概率格式良好。这里的任何数字都不是模型质量的证据;贡献在于工具及其产生的负面结果,可从干净克隆中重现,除非标记为“不可重现”。

英文摘要

Triage decides who is seen first. Building an urgency classifier for patient-voice Kinyarwanda, we found our specification could be met without producing anything it was meant to secure. We report that, and the instruments that detect it, instead of a classifier. Designed for the four languages a Rwandan health centre receives, with every instrument per-language: sentences are authored in all four arms and rows generate in one, because the frame slots those three need do not exist. Our requirement asked for one million examples; generation produced them in 130 seconds. It fails four of the nine quality gates in that specification, six when rows are attributed to their source sentence. The binding gate counts distinct authored seed phrases, not rows: at our 165, no corpus passes at any row count. Two shortfalls follow and differ: 2,835 further sentences to pass the seed-count gate, 19,835 to reach the stated million rows, because a separate gate caps a seed at 50 rows. Rows come from a machine at 7,700 per second; seeds from clinicians. A row count constrains the cheap quantity, leaves the expensive one free, and so does not constrain quality at all. Two further negatives follow. An evaluation set of 17,942 rows built from nine distinct sentences supports no verdict: our gate, which counts distinct sentences, refuses 38 of its cells and reports nothing. A model trained on a corpus of uniform surface form changes its predicted urgency for 31.5% of inputs under capitalisation and 21.0% under a single typo, reported as measurement and not attribution. A model directory shipped without its tokenizer loads without error and answers its class prior on input it cannot read, with well-formed probabilities. No figure here is evidence of model quality; the contribution is the apparatus and the negative results it produced, reproducible from a clean clone except where marked NOT REPRODUCIBLE.

Comments34 pages, 10 tables, 1 figure. Negative-results and methodology paper; no model performance claims are made

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑