arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21462cs.CLcs.AIcs.LG

CyrillicQA:语音编码的秘密语言对大语言模型性能的影响

CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance

  • Humboldt-Universität zu Berlin(柏林洪堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Erik Thureck

AI总结:

该研究探讨大语言模型在处理拉丁字母标准语言时表现优异但对其他语言变体存在劣势的背景下,能否像人类一样解码语音编码的濒危语言,以评估其保护濒危语言的能力。

AI中文摘要:

由于训练数据的选择,大语言模型(LLM)在使用拉丁字母、使用者众多的语言的标准文本输入上表现最佳,而对其他语言变体则存在劣势。不过,它们也可以成为保护这类濒危语言的多功能工具。但它们是否具备与人类相同的解码语音编码语言所需的创造力和抽象能力呢?

英文摘要:

Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while disadvantaging other language varieties. Nevertheless, they can also be a versatile tool for preserving precisely such endangered languages. But do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do?

补充信息

↑