发表机构
orinu Inc.(奥里努公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究评估先进语言模型在韩语-盲文翻译上的表现,发现其存在缺陷。通过人工标注数据集测试,发现输出差且不稳定。对比发现,对小模型进行监督微调效果更好,揭示了当前语言模型的局限,证明特定监督的有效性。
AI 中文摘要
大语言模型(LLMs)在许多语言任务上表现出色,但在诸如盲文这种结构受限且对无障碍性至关重要的模态中的能力仍不明确。我们使用人工标注数据集评估了当前最先进的LLMs在双向韩语-盲文翻译上的表现。尽管期望多语言、指令微调模型能通过文本表示推广到盲文,但结果输出一直很差且不稳定,与人类判断有很大分歧。这些结果表明缺少盲文感知分词以及韩语和盲文模式之间的对齐较弱。相比之下,在相同数据上对小模型(T5-small)进行监督微调,在标准指标上比零样本和提示LLM基线有大幅且稳定的提升。我们的发现揭示了当前LLMs的系统性局限性,并证明了适度的任务特定监督的有效性。
英文摘要
Large Language Models (LLMs) perform strongly on many language tasks, but their capability in structurally constrained, accessibility-critical modalities such as Braille remains unclear. We evaluate state-of-the-art LLMs on bidirectional Korean-Braille translation using a human-annotated dataset. Despite expectations that multilingual, instruction-tuned models can generalize to Braille via text representations, we find consistently poor, unstable outputs and substantial disagreement with human judgments. These results point to missing Braille-aware tokenization and weak alignment between Korean and Braille patterns. In contrast, supervised fine-tuning of a small model (T5-small) on the same data yields large and stable gains over zero-shot and prompted LLM baselines across standard metrics (SacreBLEU, ChrF++, CER, BLEU, ROUGE-L, METEOR, CIDEr). Our findings reveal a systematic limitation of current LLMs and demonstrate the effectiveness of modest task-specific supervision.
CommentsAccepted at the LTEDI Workshop at ACL 2026