发表机构
Kazan Federal University; Tajik National University(喀山联邦大学; 塔吉克国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对低资源语言塔吉克语,提出结合LLMs与经典词典方法的电子解释词典概念架构,为相关应用NLP任务提供核心资源,适用于计算语言学等领域从业者。
AI 中文摘要
本文提出了一种利用大语言模型(LLMs)开发塔吉克语电子解释词典的概念框架。该工作的相关性源于塔吉克语缺乏功能可与高资源语言词典相媲美的综合数字词典资源,以及现代自然语言处理技术对低资源语言系统的适配有限。在对现有语言学、统计和语料库资源进行系统调研的基础上,我们提出了一种词典架构,该架构整合了形态分析、词形还原、语义聚类以及利用LLMs生成词典条目等模块。针对塔吉克语形态的黏着特性、高形态变异性,我们论证了选择子词分词的合理性,同时采用了适用于有限标注数据的参数高效微调(PEFT)策略。本研究的新颖之处在于,提出了首个塔吉克语解释词典的整体概念架构,将经典词典编纂方法、语言统计以及LLMs的生成能力统一到单一系统中。该研究的实践意义在于,为开发功能齐全的电子词典奠定了方法论基础,该词典既可作为词典编纂工具,又可作为机器翻译、自动摘要、情感分析及其他应用自然语言处理任务的核心资源。本文面向计算语言学、词典编纂领域的专家,以及从事低资源语言相关自然语言处理系统开发的人员。
英文摘要
This paper presents a conceptual framework for developing an electronic explanatory dictionary of the Tajik language using large language models (LLMs). The relevance of the work stems from the absence of a comprehensive digital lexicographic resource for Tajik that is comparable in functionality to dictionaries for high-resource languages, and from the limited adaptation of modern natural language processing technologies to low-resource language systems. Based on a systematic survey of existing linguistic, statistical, and corpus resources, we propose a dictionary architecture that integrates modules for morphological analysis, lemmatization, semantic clustering, and dictionary entry generation using LLMs. The choice of subword tokenization is justified by the agglutinative nature of Tajik morphology and its high morphological variability, along with a parameter-efficient fine-tuning (PEFT) strategy suitable for limited annotated data. The novelty of the work lies in proposing the first holistic conceptual architecture of an explanatory dictionary for Tajik that unifies classical lexicographic methods, language statistics, and generative capabilities of LLMs into a single system. The practical significance of the study is the formation of a methodological foundation for developing a full-featured electronic dictionary that can serve both as a lexicographic tool and as a core resource for machine translation, automatic summarization, sentiment analysis, and other applied NLP tasks. The paper is intended for specialists in computational linguistics, lexicography, and developers of natural language processing systems working with low-resource languages.
Comments16 pages, 3 figures, 1 table. Preprint