发表机构
Universidad Autónoma de Madrid(马德里自治大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究介绍了可监测西班牙媒体英语借词的自填充数据库Observatorio Lazaro,其通过神经序列标注模型检测借词,经评估精度较高,能为借词研究提供持续更新的资源。
AI 中文摘要
本文介绍了Observatorio Lázaro,这是一种监测西班牙数字媒体中未同化词汇借词(主要是英语词汇借词即anglicisms)的语言资源。自2020年4月以来,该系统已自动处理一系列新闻媒体的每日内容,使用神经序列标注模型检测借词,并通过公共网页界面和API提供结果。由此形成一个持续更新的历时数据库,撰写本文时,该数据库已记录2020-2026年间188万篇文章、9.93亿个文本运行标记中的超过200万条借词。本文对该资源进行了说明:描述了端到端流程(获取、检测、后处理、存储和访问)、数据模型和可用性条款;通过检测器的保留性能(借词类的跨度级F1值为0.86)、训练语料库的标注者间一致性(Cohen's kappa值为0.91)以及对部署数据中1000个跨度的手动精度审核对该资源进行评估;并将其与西班牙语借词词典学、带标注的借词语料库和新词监测观测站进行对比。数据显示,未同化的anglicisms在西班牙媒体中的使用频率约为每千个标记2个,且该频率保持稳定。我们对六年的统计分析表明,西班牙语中的anglicisms词汇表现为开放且不断增长的类别,其中58.7%的类型仅出现一次(校正检测精度后为53.6%),其密度在时尚、科技和生活方式板块最高,在政治和机构新闻板块最低。该资源旨在补充静态借词词典和一次性带标注语料库,提供西班牙媒体中借词的持续更新记录。
英文摘要
This paper describes Observatorio Lázaro, a language resource that monitors unassimilated lexical borrowings (predominantly English lexical borrowings or anglicisms) in the Spanish digital press. Since April 2020 the system has automatically processed the daily output of a collection of news outlets, detected borrowings with a neural sequence-labeling model, and made the results available through a public web interface and API. The result is a continuously updated diachronic database which, at the time of writing, records more than two million borrowings across 1.88 million articles and 993 million running tokens of text (2020-2026). The paper documents the resource: we describe the end-to-end pipeline (acquisition, detection, post-processing, storage and access), the data model and the terms of availability; we evaluate the resource through the detector's held-out performance (span-level F1=0.86 for the borrowing class), inter-annotator agreement on the training corpus (Cohen's kappa=0.91) and a manual precision audit of 1,000 spans from the deployed data; and we situate it with respect to Spanish borrowing lexicography, annotated borrowing corpora and neology-monitoring observatories. The data shows that unassimilated anglicisms are used in the Spanish press at a frequency of approximately two anglicisms per thousand tokens, and that this rate remains stable. Our statistical analysis over six years reveals that the anglicism vocabulary in Spanish behaves as an open and growing class, with 58.7% of its types attested only once (53.6% after correcting for detection precision), and that its density is highest in the fashion, technology and lifestyle sections and lowest in political and institutional news. The resource is intended to complement static borrowing dictionaries and one-off annotated corpora by providing a continuously updated record of borrowing in the Spanish press.