arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

标本馆标签数字化自动化流程

Automated pipeline for herbarium label digitization

Hiba Abbad, Hanane Ariouat, Eva Perez Pimpare, Nicolas Turenne, Eric Chenin, Abderrazak Sebaa, Edi Prifti, Jean-Daniel Zucker, Youcef Sklab

arXiv 2608.28676首次发表:更新:

发表机构

École Supérieure en Sciences et Technologies de l’Informatique et du Numérique (ESTIN); IRD; Sorbonne Université; Muséum national d’histoire naturelle; INRAE; INSERM; AP-HP(高等计算机与数字科学技术学院(ESTIN); 法国发展研究所; 索邦大学; 法国国家自然历史博物馆; 法国国家农业食品与环境研究院; 法国国家健康与医学研究院; 巴黎公共医疗集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出HERBIOME自动化流程,整合YOLOv8、CRAFT Hezar、TrOCR、GPT-4o Mini等技术,实现标本馆标签的元数据结构化提取,构建图像-文本数据集,推动多模态生物多样性AI发展。

AI 中文摘要

数字化标本馆藏品目前已包含超过1亿张可免费访问的标本图像,成为解决生态学与进化生物学基础问题的关键资源。然而,标本馆标签中编码的丰富元数据(采集者身份、地理地点、采集日期及生态观测结果)仍无法大规模获取,这既限制了生物多样性信息学的发展,也阻碍了用于多模态AI的标本专属图像-文本语料库的构建。本文提出HERBIOME,这是一个模块化端到端的自动化标本馆标签数字化流程,整合了基于YOLOv8的组件检测、CRAFT Hezar词级文本定位、微调后的TrOCR(用于识别混合手写与印刷文本),以及GPT-4o Mini(用于将语义元数据结构化到标准化字段)。TrOCR在多源数据集上进行训练,该数据集结合了通用转录语料库(CREMMA-AN、PictoCatalogs)与标本馆专属数据(RéColNat),字符错误率达到4.05%-4.10%。对450份法国标本馆标本进行端到端评估,采用最大窗口相似度(MWS:0.614-0.618)与语义元数据准确率(SMA:0.440-0.445)的双指标框架,结果显示混合训练策略可提升语义保真度,而随机采样可最大化表面相似度,分类学字段仍是主要瓶颈。HERBIOME通过从复杂异构标签中自动提取结构化元数据,减少了转录负担,能够构建忠实捕捉标本个体特征的配对图像-文本数据集,这是下一代多模态生物多样性AI系统的前提条件。

英文摘要

Digitized herbarium collections, now comprising over 100 million freely accessible specimen images, have become a critical resource for addressing fundamental questions in ecology and evolutionary biology. Yet the rich metadata encoded in herbarium labels (collector identities, geographic localities, collection dates, and ecological observations) remains largely inaccessible at scale, constraining both biodiversity informatics and the construction of specimen-specific image-text corpora for multimodal AI. We present HERBIOME, a modular end-to-end pipeline for automated herbarium label digitization, integrating YOLOv8-based component detection, CRAFT Hezar word-level text localization, fine-tuned TrOCR for recognition of mixed handwritten and printed text, and GPT-4o Mini for semantic metadata structuring into standardized fields. TrOCR was trained on a multi-source dataset combining general transcription corpora (CREMMA-AN, PictoCatalogs) with herbarium-specific data (RéColNat), achieving a Character Error Rate of 4.05-4.10%. End-to-end evaluation on 450 French herbarium specimens, using a dual-metric framework of Maximum Window Similarity (MWS: 0.614-0.618) and Semantic Metadata Accuracy (SMA: 0.440-0.445), reveals that hybrid training strategies improve semantic fidelity while random sampling maximizes surface similarity, with taxonomic fields remaining the principal bottleneck. By automating the extraction of structured metadata from complex, heterogeneous labels, HERBIOME reduces transcription burden, enables the construction of paired image-text datasets that faithfully capture specimen individuality, which is a prerequisite for next-generation multimodal biodiversity AI systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑