arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

非完美标注下牛科动物牙列的分割:卷积模型与注意力模型的对比研究

Segmentation of Bovid Dentition Under Imperfect Annotations: A Comparative Study of Convolutional and Attention Models

Keith G. Mills, Evan B. Sanders, Gregory J. Matthews, Juliet K. Brophy

arXiv 2608.31052首次发表:更新:

发表机构

Episcopal High School of Baton Rouge(巴吞鲁日圣公会高中)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对B.O.V.I.D.数据集的非完美标注,对比卷积模型与注意力模型的牙列分割性能,发现预处理技术对定量指标影响有限但定性影响显著。

AI 中文摘要

语义分割将图像分解为对应不同对象类别的不同掩码区域,如人、汽车、标志或建筑物。机器学习(ML)的进展使该任务从边缘检测等传统基于规则的启发式方法转向深度学习(DNN),后者可直接学习对像素进行分类。然而,语义分割DNN关键依赖专家设计的掩码目标来学习,不完美或未对齐的掩码会干扰模型的有效学习能力。本文对应用于B.O.V.I.D.数据集(包含高分辨率牛科动物牙列照片及专为ML训练设计的手工分割掩码)的分割架构进行对比研究,范围涵盖卷积骨干网络到视觉Transformer。我们评估一系列预处理和对齐技术以缓解标签缺陷,发现这些预处理选择对Dice分数和mIoU等定量指标的影响有限,但对预测掩码的定性影响显著。

英文摘要

Semantic segmentation decomposes an image into distinct mask regions corresponding to different object categories, such as people, cars, signs or buildings. Advances in machine learning (ML) have shifted this task away from traditional rule-based heuristics such as edge detection, towards deep neural networks (DNN) that learn to classify pixels directly. However, semantic segmentation DNNs crucially depend on expertly designed mask targets to learn from, and imperfect or misaligned masks can interfere with a model's ability to learn effectively. This paper presents a comparative study of segmentation architectures, ranging from convolutional backbones to vision transformers, applied to the B.O.V.I.D. dataset, a corpus of high-resolution bovid dental photographs paired with hand-made segmentation masks not originally designed for ML-based training. We evaluate a range of preprocessing and alignment techniques to mitigate the resulting label imperfections. We find that while these preprocessing choices have limited effect on quantitative metrics such as Dice score and mIoU, their qualitative impact on predicted masks is substantial.

Comments13 pages, 2 Tables, 13 Figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑