发表机构
College of Science and Technology, Royal University of Bhutan(不丹皇家大学科学技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对宗卡语记录难、打字不便的问题,通过减少击键次数让其打字更便捷。利用DCDD数据集经预处理后,用LSTM、Bi-LSTM和GRU三种模型训练,GRU表现最佳,准确率74.03%且解决过拟合问题。
AI 中文摘要
宗卡语是不丹的国语,该国的官方文件、经文等都用宗卡语书写以保留文化价值。但记录宗卡语写作具有挑战性且耗时,因为其文字复杂、每个音节需多次击键且高效打字工具有限。该项目旨在通过减少击键次数让宗卡语打字更便捷。数据集来自DCDD,经数据预处理后,用LSTM、Bi-LSTM和GRU三种模型训练,微调超参数。GRU表现最佳,准确率达74.03%,还解决了过拟合问题。
英文摘要
Dzongkha, being the national language of Bhutan, is a common and widely spoken language in the country. Official documents, scriptures and other literature products are written in Dzongkha in order to retain the cultural value. However, documenting Dzongkha writing is a challenging and time-consuming process, largely due to the complexity of the script, the need for multiple keystrokes per syllable, and the limited availability of efficient typing tools. An immediate system that can predict and display a list of probable words for Dzongkha is the solution for this problem. The project is mainly aimed to make Dzongkha typing as convenient as possible by reducing the number of keystrokes. Our dataset is acquired from DCDD and has a total of 100000 sentences, 1331282 words and 28344 unique words. The data preprocessing was done by removing all the alphanumeric characters, tokenization, generating N-gram sequences and padding. Three models selected for training are LSTM, Bi-LSTM and GRU. The training process included fine-tuning of the model's hyperparameters. GRU being lightweight and able to handle larger datasets performed best with 74.03% accuracy and also solved the problem of overfitting.
Comments6 pages, 7 figures, 4 tables