arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

约鲁巴语在Unicode中的问题概述

Yorùbá in Unicode: An Overview of a Problem

Kólá Túbòsún

arXiv 2609.33734首次发表:更新:

发表机构

Yoruba Names Project(约鲁巴名字项目)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文概述约鲁巴语在Unicode中因缺乏预组合字符而导致的书写与搜索问题,指出NFC规范化策略为结构性障碍,并提议正式编码四个核心字符以持久解决。

AI 中文摘要

在互联网和计算机上书写约鲁巴语时,存在一个长期反复出现且多年来难以解决的问题。该语言以及依赖变音符号进行消歧的其他非洲语言,需要一小部分Unicode未编码的预组合字符。这迫使作者和数字系统依赖组合字符序列,这些序列在不同平台上表现不一致,在字体替换时损坏,并在搜索中失败。本文通过个人和经验证据,记录了从已出版书籍到网络平台再到移动键盘的一系列真实世界情境中的这种失败。它指出Unicode的NFC规范化稳定性策略是阻碍直接修复的结构性约束,主张联盟直接干预以解决这一活跃问题,并提出对四个核心约鲁巴语字符的正式编码请求作为最持久的解决路径。

英文摘要

There is a recurrent problem in the writing of Yorùbá on the internet and on the computer that has proven intractable over the years. The language, along with other African languages that depend on diacritics for disambiguation, requires a small set of precomposed characters that Unicode does not encode. This has forced writers and digital systems to rely on combining character sequences that behave inconsistently across platforms, corrupt under font substitution, and fail in search. This paper documents that failure across a range of real world contexts, from published books to web platforms to mobile keyboards, using personal and empirical evidence. It identifies Unicode's NFC normalization stability policy as the structural constraint that prevents a straightforward fix, arguing for direct intervention of the Consortium in solving the active problem, proposing a formal encoding request for the four core Yorùbá characters as the most durable path to resolution.

Commentsv3 corrected explanation of why certain vowels lack precomposed forms; code-point typo fixed

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑