发表机构
Alliance School of Liberal Arts and Sciences; Alliance University; Department of Computer Science and Engineering(联盟文理学院; 联盟大学; 计算机科学与工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过样本复杂度曲线直接测量手写天城文识别中预训练对标注成本的影响,发现监督式合成预训练可大幅减少所需真实转录单词数,但优势随目标精度提高而减弱。
AI 中文摘要
训练手写文本识别系统需要单词图像及其对应的转录文本,而这些转录文本是人工生成的。对于一种只有少数专家能够阅读的文字,这种人工转录成为一个限制因素,因为训练出的模型本应节省这些专家的时间。因此,一个相关问题浮现出来:在识别器变得有用之前需要多少转录文本,以及预训练能消除多少这种成本?在本研究中,我们直接针对手写天城文测量了答案。我们保持识别器、优化器和评估协议不变,仅改变用于微调的真实转录单词数量,涵盖从10到4000的九个预算,以及四种初始化机制,每个数据点使用六个随机种子。得到的曲线随后被转换为标注等效项。通过监督式合成预训练,仅使用81个转录单词即可达到0.50的字符错误率,而随机初始化需要355个,这给出了4.40 [3.56, 4.99]的标签乘数。还存在一个零样本参考点:完全没有真实转录单词时,这种预训练的价值相当于约136个真实转录单词。随着目标准确率的提高,这一优势逐渐减小,在我们测量的最严格目标下,该优势与无节省无法区分。第四个分支仅迁移编码器,将预训练方法的效果与迁移范围的效果分开,并观察到掩码图像建模在有限的预算范围内产生负迁移。我们强调,本研究中的稀缺性是通过对大型语料库进行子采样构建的。
英文摘要
To train handwritten text recognition systems we need word images and their corresponding transcriptions, and these transcriptions are produced manually. For a script that can be read by only a small number of specialists, this manual transcription is a limitation, because the trained models are supposed to save the time of those same specialists. A relevant question therefore arises: how many transcriptions are needed before a recogniser becomes useful, and how much of that cost can pretraining remove? In this study the answer is measured directly for handwritten Devanagari. We keep the recogniser, optimiser and evaluation protocol the same and change only the number of real transcribed words used for fine-tuning across nine budgets from 10 to 4,000 and four initialisation regimes, with six seeds at every point. The resulting curves are then converted into annotation-equivalent terms. A CER of 0.50 is reached by supervised synthetic pretraining using only 81 transcribed words, whereas random initialisation requires 355, which gives a label multiplier of 4.40 [3.56, 4.99]. There is a zero-shot reference point as well: with no real transcribed words at all, this pretraining is worth about 136 of them. This advantage gets smaller as the target accuracy improves, and at the most demanding target we measure, it cannot be distinguished from no saving at all. A fourth arm in which only the encoder is transferred separates the effect of the pretraining method from that of transfer scope, and masked image modelling is observed to transfer negatively over a bounded range of budgets. We emphasise that the scarcity in this study is constructed by subsampling a large corpus.