AI 中文总结
研究针对现有多语言自然语言处理度量基于Unicode码点操作的问题,引入字素工具包,将度量扩展到字素簇操作,还改进了特定语言的字素处理,通过案例研究证明该工具包能更准确评估复杂脚本。
AI 中文摘要
现有词汇距离、相似度和评估度量基于Unicode码点操作,这在单个字素由多个Unicode码点表示的书写系统中可能误判错误。我们引入了字素工具包,一个开源Python库,将这些度量扩展到基于字素簇操作。该库还为泰米尔语和僧伽罗语提供了改进的字素处理,包括准确的字素簇识别和字素合成/分解实用工具。通过一个光学字符识别案例研究,我们证明字素级度量能更如实地评估复杂脚本。
英文摘要
Existing lexical distance, similarity, and evaluation metrics operate on Unicode code points, which can misrepresent errors in writing systems where a single grapheme is represented by multiple Unicode code points. We introduce grapheme-kit, an open-source Python library that extends these metrics to operate on grapheme clusters instead. The library also provides improved grapheme processing for Tamil and Sinhala, including accurate grapheme cluster identification and grapheme composition/decomposition utilities. Through an OCR case study, we demonstrate that grapheme-level metrics provide a more faithful evaluation of complex scripts.