绘制亚美尼亚巴黎:从20世纪侨民报刊中提取并地理编码商业广告
Mapping Armenian Paris: Extracting and Geocoding Commercial Advertisements from the 20th-Century Diaspora Press
浏览论文内容
中文总结 AI 辅助
本研究提出基于IIIF的VLM驱动端到端流程,从西亚美尼亚语报刊中提取并地理编码商业广告,构建巴黎亚美尼亚商业社区交互式地图,相关成果可迁移至其他资源匮乏历史语料库。
中文摘要 AI 辅助
本文提出一种基于IIIF的端到端流程,将法国的数字化亚美尼亚报刊转化为20世纪巴黎亚美尼亚商业社区的交互式地图。流程会定位每一页的商业广告,对其进行读取并解析为结构化记录,随后将这些记录地理编码并放置在地图上。西亚美尼亚语资源匮乏,且缺乏现成的布局和OCR模型支持,因此该流程采用视觉语言模型(VLM)作为数据自举策略:这类模型能生成人工标注无法达到的规模的可用结构化记录,且在常规行级CRNN OCR失效的强弯曲扫描件上仍保持可靠。研究贡献包括含3270条广告级标注的500页西亚美尼亚语报刊语料库、可在单次标注过程中捕获检测与语义字段的Label Studio模板,以及可迁移至其他资源匮乏历史语料库的可复现工作流。更广泛而言,该研究表明VLM驱动的数据自举是针对西亚美尼亚语这类资源匮乏的历史语言的有效手段。
英文摘要
This paper presents an end-to-end, IIIF-based pipeline that turns the digitised Armenian press of France into an interactive map of the 20th-century Parisian Armenian commercial community. On each page, commercial advertisements are located, read, and parsed into structured records, which are then geocoded and placed on the map. Western Armenian is under-resourced and unsupported by off-the-shelf layout and OCR models, so the pipeline uses vision-language models (VLMs) as a data-bootstrapping strategy: they produce usable structured records at a scale hand annotation could not reach, and stay reliable on the strongly curved scans where conventional line-level CRNN OCR breaks down. The contribution includes a 500-page Western Armenian press corpus with 3,270 advertisement-level annotations, a Label Studio template that captures detection and semantic fields in a single annotation pass, and a reproducible workflow transposable to other under-resourced historical corpora. More broadly, the work shows that VLM-driven data bootstrapping is an effective lever for under-resourced historical languages such as (Western) Armenian.
发表机构
- Calfa(卡尔法机构)
- École nationale des chartes, PSL University(巴黎文理研究大学国家文献与遗产学院)
- Université française en Arménie (UFAR)(亚美尼亚法国大学)
机构由 AI 辅助整理,请以论文原文为准。