arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

韩国期刊摘要中的LLM相关语域转变:一项基于形态感知的超额词汇研究,2018-2026

An LLM-Associated Register Shift in Korean Journal Abstracts: A Morphology-Aware Excess-Vocabulary Study, 2018-2026

Aron Lee

arXiv 2609.07447首次发表:更新:

发表机构

INTFRAME Research(INTFRAME研究公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过形态感知的超额词汇分析,发现韩国KCI期刊摘要自2024年末出现LLM相关语域转变,并提出条件下限估计,表明LLM处理比例逐年上升。

AI 中文摘要

超额词汇,即某个词在2023年前趋势之上的频率,是衡量2022年后学术英语变化的方式。我们将其适配到韩语,使用398,296篇KCI摘要(2018年至2026年8月)中的形态单位,并以47,165篇越南语摘要作为对比。安慰剂下限为单词统计量的0.1-2.2个百分点,重新选择的半分集统计量最多为2.9。韩语摘要显示2023年无变化,2024年末开始出现,2025年上升,2026年中趋于平稳:sisahada“suggest”出现在2026年摘要的21.4%中,而预期为5.3%;普通动词如araboda“look into”降至趋势的四分之一。在所述假设下,2024-2026年LLM处理摘要的单词条件下限分别为3.5%、10.5%和16.1%,半分集界限为7.8%、20.6%和33.0%。Holzwarth等人的估计器在相同约束下给出2025-2026年的41.9%和72.1%。主题控制减少但未消除该现象:将集合限制为三个语言模型标注者均称为风格的词元,33.0个百分点中留下14.7个;将每篇2026年摘要与其期刊最接近的基期摘要配对,留下34.1个。测试的翻译路径无法解释该现象:随着标记上升,翻译韩语的表面标记下降。在同一篇文章的英文摘要中,超额出现早一年;在英文侧无此现象的地方,韩语转变持续存在,其速率相当于有现象处的30%至66%。来自三个提供者的对照摘要再现了上升词汇,标记更替与模型生成一致;隐含流行率取决于情景。

英文摘要

Excess vocabulary, a word's frequency above its pre-2023 trend, is how the change in scholarly English after 2022 has been measured. We adapt it to Korean with morphological units on 398,296 KCI abstracts (2018-August 2026), with 47,165 Vietnamese abstracts for comparison. Placebo floors are 0.1-2.2 points for the single-word statistic and at most 2.9 for the re-selected split-half set statistic. Korean abstracts show nothing in 2023, onset in late 2024, a rise through 2025 flattening in mid-2026: sisahada "suggest" appears in 21.4% of 2026 abstracts against 5.3% expected; plain verbs like araboda "look into" fall to a quarter of trend. Under stated assumptions the single-word conditional lower bound on LLM-processed abstracts is 3.5%, 10.5% and 16.1% for 2024-2026 and a split-half set bound 7.8%, 20.6% and 33.0%. Holzwarth et al.'s estimator under the same discipline gives 41.9% and 72.1% for 2025-2026. Subject-matter controls reduce but do not remove it: restricting the set to lemmas three language-model annotators all call style leaves 14.7 of the 33.0 points, and pairing each 2026 abstract with its journal's closest base-period abstract leaves 34.1. Tested translation routes do not explain it: the surface marks of translated Korean fall as the markers rise. In the same articles' English abstracts the excess appears a year earlier; where the English side carries none, the Korean shift persists at 30 to 66% of the rate where it does. Control abstracts from three providers reproduce the rising words, with marker turnover consistent with model generations; implied prevalences are scenario-dependent.

Comments30 pages, 8 figures, 31 tables. Working paper. Data, code and this version are deposited at Zenodo: doi:10.5281/zenodo.22303588

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑