arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从转录到语义语料库分析:古代语言句子表示的无监督学习

From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages

Th{é}otime de la Selle

arXiv 2607.24542首次发表:更新:

AI 中文总结

研究古代语言句子表示的无监督学习,采用TSDAE和CSE策略,将专门语言模型转换为特定语料库句子编码器,在教父文献圣经重用案例中表现出色,优于多种基线,且可跨作品作者迁移,相关工具可供非专业人员使用。

AI 中文摘要

自动文本识别(ATR)为数字人文提供了大量古代语言的非结构化、异构且通常有噪声的文本。下游语义分析(文本重用识别、对齐和语义搜索)依赖句子嵌入,但现有方法难以应用于古代语言。我们研究了两种完全无监督的策略——TSDAE和对比句子嵌入(CSE),它们仅使用原始句子将专门的词元级语言模型转换为特定语料库的句子编码器。在教父文献中圣经重用的文献学核心案例上,我们将重用识别分解为两个单独评估的任务——二元检测和对应检索,并针对多语言、专门、蒸馏和监督微调的基线以及模拟HTR工件和抄写缩写的人工噪声数据对适配的编码器进行基准测试。适配的编码器在两项任务上均优于所有基线,具有互补的特点:TSDAE在有大量领域内语料库时引领检测,而CSE引领检索,在仅4 - 8k个原始领域内句子时达到最优——在笔记本电脑GPU上只需几十秒的训练——并且可以跨作品和作者进行迁移,包括直接在有噪声的ATR后文本上重新训练时。UMAP地图集将每种策略的几何效应与测量的增益相关联,完整的管道(分割、微调、跨语料库语义搜索)通过在线工具Paraphrasis提供给非专业人员。

英文摘要

Automatic Text Recognition (ATR) now supplies digital humanities with large volumes of unstructured, heterogeneous, and often noisy text in ancient languages. Downstream semantic analysestext reuse identification, alignment, and semantic search-rely on sentence embeddings, yet existing methods transfer poorly to ancient languages: generic multilingual encoders underperform, specialized language models yield anisotropic representation spaces, and labeled similarity data is unavailable. We study two fully unsupervised strategies - TSDAE and contrastive sentence embedding (CSE) - that adapt a specialized token-level language model into a corpus-specific sentence encoder using only raw sentences. On the philologically central case of biblical reuse in patristic literature (2,935 expert-verified parallels in Latin and Ancient Greek, from Augustine, Jerome, and Athanasius), we decompose reuse identification into two separately evaluated tasks-binary detection and correspondence retrieval-and benchmark the adapted encoders against multilingual, specialized, distilled, and supervised fine-tuned baselines, as well as on artificially noised data simulating HTR artifacts and scribal abbreviations. The adapted encoders outperform all baselines on both tasks, with complementary profiles: TSDAE leads detection given a large in-domain corpus, while CSE leads retrieval, reaches its optimum with as few as 4-8k raw in-domain sentences-a few tens of seconds of training on a laptop GPU-and transfers across works and authors, including to noisy post-ATR text when retrained directly on it. UMAP atlases relate the geometric effect of each strategy to the measured gains, and the full pipeline-segmentation, fine-tuning, cross-corpus semantic search-is made available to non-specialists through the online tool Paraphrasis.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑