arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

pylazaro:一个用于西班牙语中英语外来词提取的Python包

pylazaro: a Python package for anglicism extraction in Spanish

Elena Alvarez-Mellado

arXiv 2609.29276首次发表:更新:

发表机构

Universidad Autónoma de Madrid(马德里自治大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文介绍pylazaro,一个用于从西班牙语文本中自动提取英语外来词的开源Python包,提供统一接口集成五个序列标注模型,性能优于通用LLM(F1 0.86 vs 0.40),已被下载超5.8万次并用于监测西班牙媒体中的英语外来词。

AI 中文摘要

词汇借用是指一种语言中的词汇被引入到另一种语言中。在文本中识别词汇借用对于语言学中的数据密集型领域(如词典编纂或语料库语言学)是一项相关任务,但标准的文本处理库中没有一个提供这样的功能。在本文中,我们介绍了pylazaro,一个用于从西班牙语文本中自动提取未同化的词汇借用(主要是英语外来词)的开源Python包。pylazaro为五个序列标注模型提供了统一接口,这些模型使用不同的库进行训练,因此用户可以运行和切换它们,而无需处理每个库的特殊性。我们描述了该包的设计和用法,将其模型的性能与通用大语言模型(LLM)进行了对比(后者在此任务上表现不佳:F1分数低于0.40,而pylazaro中最佳模型的F1分数为0.86),并报告了其采用情况:pylazaro已被下载超过58,000次,并且是Observatorio Lazaro背后的库,该资源用于监测西班牙媒体中的英语外来词使用情况。pylazaro可以通过PyPI安装,在readthedocs中有文档,并可通过托管在HuggingFace Spaces上的实时演示进行试用。

英文摘要

Lexical borrowings are words from one language that are introduced into another language. Identifying lexical borrowings in text is a relevant task for data-centric fields in Linguistics such as lexicography or corpus linguistics, but none of the standard libraries for text processing offers such a functionality. In this paper we present pylazaro, an open-source Python package for the automatic extraction of unassimilated lexical borrowings (mostly anglicisms) from Spanish text. pylazaro offers a single interface to five sequence labeling models that were trained using different libraries, so that users can run and switch between them without having to deal with the idiosyncrasies of each library. We describe the design and usage of the package, contrast the performance of its models with that of general-purpose LLMs (which perform poorly at this task: F1 below 0.40, compared to 0.86 for the best model in pylazaro) and report on its adoption: pylazaro has been downloaded more than 58,000 times and is the library behind Observatorio Lazaro, a resource that monitors anglicism usage in the Spanish press. pylazaro can be installed via PyPI, is documented in readthedocs and can be tried through a live demo hosted on HuggingFace Spaces.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑