arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于鸟类骨骼分类的视觉与形态计量特征多模态融合

Multimodal fusion of visual and morphometric features for avian bone classification

Nevio Dubbini, Lisa Yeomans, Marco Pavia, Ramazan Parmaksiz, Ayse Atas Hooglugt, Gabriele Gattiglia, Beatrice Demarchi

arXiv 2607.26743首次发表:更新:

AI 中文总结

本研究提出结合卷积神经网络图像分析与骨计量测量的多模态框架,用于鸟类骨骼分类,在骨骼类型分类准确率达86%,科级分类Top-1准确率51%、Top-3准确率75%,为AI辅助动物考古识别建立了方法学基线。

AI 中文摘要

人工智能在考古应用中已展现出巨大潜力,但在动物考古学领域的应用仍有限,尤其在鸟类骨骼遗骸的识别方面。本研究提出一种概念验证多模态框架,将基于卷积神经网络的图像分析与骨计量测量相结合,用于鸟类骨骼分类。使用来自多个博物馆和研究机构数据集的10000余张图像,研究了两项分类任务:骨骼部位识别和科级分类。分类前,采用结合BiRefNet与SAM2的两阶段流程自动对图像进行分割。利用预训练的EfficientNet_V2_S骨干网络提取的视觉特征,通过特征级多模态架构与标准化形态计量数据融合。该模型在测试集的骨骼类型分类任务中准确率达86%,证明对骨骼部位的可靠识别;科级分类更具挑战性,Top-1准确率为51%但Top-3准确率达75%,表明正确分类单元常出现在最可能的预测结果中。这些结果证明了在统一深度学习框架内结合视觉与形态计量信息的可行性,并为未来AI辅助动物考古识别建立了方法学基线,该方法助力开发可扩展、可解释且符合考古学意义的鸟类遗骸研究工具。

英文摘要

Artificial intelligence has shown considerable potential for archaeological applications, yet its use in zooarchaeology remains limited, particularly for the identification of avian skeletal remains. This study presents a proof-of-concept multimodal framework that integrates convolutional neural network-based image analysis with osteometric measurements for the classification of bird bones. Using a dataset of more than 10,000 images from multiple museum and research collections, two classification tasks were investigated: skeletal element identification and family-level taxonomic classification. Prior to classification, images were automatically segmented using a two-stage pipeline combining BiRefNet and SAM2. Visual features extracted with a pre-trained EfficientNet_V2_S backbone were fused with standardized morphometric data through a feature-level multimodal architecture. The model achieved 86% accuracy on the test set for bone-type classification, demonstrating reliable recognition of skeletal elements. Family-level classification proved more challenging, reaching 51% top-1 accuracy but 75% top-3 accuracy, indicating that correct taxa were frequently included among the most probable predictions. These results demonstrate the feasibility of combining visual and morphometric information within a unified deep-learning framework and establish a methodological baseline for future AI-assisted zooarchaeological identification. The approach contributes to ongoing efforts to develop scalable, interpretable, and archaeologically meaningful tools for the study of avian remains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑