arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08037cs.CL

语言模型是否具备文字系统意识?

Are Language Models Script-Aware?

David Kletz, Sandra Mitrović, Ljiljana Dolamić, Fabio Rinaldi

首次发表
浏览论文内容

中文总结 AI 辅助

本文探究语言模型是否具备文字系统知识,通过两个实验发现模型能匹配输入文字系统并遵循指令,且LLMs表现优于SLMs。

中文摘要 AI 辅助

语言模型经常生成非预期语言或文字系统的输出,这一现象被称为脱靶生成。尽管现有研究聚焦于语言选择,但文字系统知识这一维度仍研究不足:在进行任何语言理解之前,用户必须识别模型响应中的图形符号。我们通过在多文字系统语言上测试小型和大型语言模型(SLMs 和 LLMs),探究它们是否具备文字系统知识。通过两个互补实验,我们评估模型是否(1)调整其输出文字系统以匹配输入,以及(2)遵循明确指令以指定文字系统生成文本。我们测试的模型展现出显著的文字系统知识:它们均实现了近乎完美的拉丁文字保真度(超过98%),并以高频率遵循文字系统指令。尽管如此,我们注意到LLMs和SLMs之间存在差异,LLMs在包括非标准文字系统组合在内的得分更高。

英文摘要

Language models frequently generate outputs in unintended languages or scripts, a phenomenon known as off-target generation. While existing research has focused on language selection, the dimension of script knowledge remains understudied: before any linguistic understanding can occur, users must recognize the graphic symbols in a model's response. We investigate whether Small and Large Language Models (SLMs and LLMs) possess script knowledge by testing them on multi-scriptic languages. Through two complementary experiments, we evaluate whether models (1) adapt their output script to match the input, and (2) follow explicit instructions to generate text in a specified script. The models we tested demonstrate substantial script knowledge: they all achieve a near-perfect Latin script fidelity (more than 98%) and follow script instructions with high frequency. Nevertheless, we notice differences between LLMs and SLMs, with higher scores for LLMs including for non-standard script combinations.

发表机构

  • SUPSI, IDSIA, Switzerland(瑞士南部应用科学与艺术大学,IDSIA)
  • armasuisse, Science & Technology, Switzerland(瑞士军备局,科学与技术部)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑