arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06275cs.CVeess.IV

TLNM:使用 Mask R-CNN 从智能手机照片进行经外部验证的牙齿检测、编号与分割

TLNM: Externally Validated Tooth Detection, Numbering and Segmentation from Smartphone Photographs Using Mask R-CNN

Arash Nedaei, Henna Tiensuu, Elina Väyrynen, Saujanya Karki, Jaakko Suutala

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出基于 Mask R-CNN 的 TLNM 模型,经1272张智能手机图像训练,结合两种优化机制,在内外测试集上表现优异,可实现智能手机照片的牙齿检测、编号与分割,为远程牙科提供低成本方案。

中文摘要 AI 辅助

口腔健康问题影响全球数十亿人,但专业牙科护理的高成本与有限可及性阻碍了预防性口腔医疗。现有研究依赖临床级X光片或口腔内相机图像,这类图像无法用于公众自我筛查。本研究提出一种适用于智能手机照片的牙齿定位与编号模型。我们开发了定制化的 Mask R-CNN(掩码区域卷积神经网络)流程,该流程在1272张带标注的智能手机图像上完成训练。为应对用户生成健康数据的变异性,该流程融入两种领域知识驱动的机制:一是掩码灰度世界白平衡算法,用于减轻人工色偏;二是解剖学约束检测层,用于强化结构有效性并抑制假阳性。评估包含四个阶段:内部保留测试、独立外部测试、描述性 ablation 研究(消融研究),以及使用同一内部测试集的折基训练稳定性分析。在内部测试集上,该模型的实例掩码 AP@50 达0.818、类别感知 PQ 达0.780、操作 F1 达0.884;训练稳定性表现为模型间差异有限:十次运行中,实例掩码 AP@50 的标准差为0.009。在外部数据集上,尽管存在人群、传感器及采集方案的差异,该模型仍取得实例掩码 AP@50 达0.901、类别感知 PQ 达0.832、操作 F1 达0.928的结果。推理流程以开源容器化 API 形式提供。这些结果表明,消费级智能手机图像可支持自动化牙齿层面的解剖映射,为资源受限环境下的远程筛查与远程牙科提供了可扩展、潜在低成本的基础。

英文摘要

Oral health issues affect billions of people globally, but the cost and limited access to professional dental care hinder preventive oral healthcare. Research relies on clinical-grade sensors, unavailable for public self-screening. This study introduces a tooth localisation and numbering model for smartphone photographs. We developed a customised Mask Region-based Convolutional Neural Network pipeline trained on 1,272 annotated smartphone images. To address variability in patient-generated health data, the pipeline incorporates two domain-informed mechanisms: a masked gray-world white-balancing algorithm to mitigate artificial colour casts and an anatomically constrained detection layer to enforce structural validity and suppress false positives. The system was evaluated using internal and external testing, descriptive ablation study and training stability analysis. On the internal test set, the model achieved an instance-mask AP50 of 0.818, class-aware PQ of 0.780, and operational F1 of 0.884. On the external dataset, the model achieved an instance-mask AP50 of 0.901, class-aware PQ of 0.832, and operational F1 of 0.928 despite differences in population, sensors, and acquisition protocols. The inference pipeline is available as an open-source, containerised API. These results demonstrate that consumer-grade smartphone imagery can support automated tooth-level anatomical mapping, offering a scalable, potentially low-cost foundation for remote screening and tele-dentistry in resource-constrained environments.

发表机构

  • University of Oulu(奥卢大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑