发表机构
DiceMed; University of Maryland, College Park; Indira Gandhi National Open University; Indian Institute of Technology Jammu(DiceMed; 马里兰大学学院公园分校; 英迪拉·甘地国立开放大学; 印度理工学院贾姆穆分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种利用已配准口内数据直接测量咬合几何的方法,通过梯度提升和ConvNeXt-Tiny融合生成正畸报告,在ODIN 2026挑战中排名第三。
AI 中文摘要
从口内数据生成正畸报告通常被建模为多模态字幕生成任务,然而发布的Bite2Text扫描对在咬合状态下已经配准,这使得几个核心咬合量可以直接测量而非推断。本文报告的系统利用了这一特性:每例病例的解剖坐标系通过牙弓锥度和牙弓闭合恢复,而非采用在数据集中不一致的RAS约定,每个牙弓被简化为牙弓角度坐标下的咬合嵴轮廓,从而以闭合形式获得覆𬌗、覆盖、中线偏移、横向重叠、反𬌗范围、尖牙交错滞后以及咬合曲线。梯度提升将31个此类测量值映射到13个模板字段,仅当患者级交叉验证优于其自身的多数基线时才预测该字段,确定性渲染器生成语料库的六部分叙述;一个ConvNeXt-Tiny分类器在五个标准化摄影视图上融合每个字段,将平均字段准确率从0.601提高到0.683。对挑战评估器的重新实现表明,其BLEU-4和METEOR是局部变体,其F均值权重中召回率与精确率之比为9:1,两位临床医生对同一患者的发现一致性为47%,因此恒定报告在字幕生成上比真正的第二临床医生报告高出0.165。留出集分数在口内扫描参考上达到BLEU-4 0.458和METEOR 0.677,在照片参考上达到0.278和0.507,提交的系统在ODIN 2026 Bite2Text测试阶段以0.2680和0.4629排名第三,距第一名仅差0.022 BLEU-4,在CPU上每例运行时间不到十秒。数据集和代码可在该https URL获取。
英文摘要
Orthodontic report generation from intraoral data is normally cast as multimodal captioning, yet the released Bite2Text scan pairs are supplied already registered in occlusion, which makes several core occlusal quantities directly measurable rather than inferable. The system reported here exploits that property: an anatomical frame is recovered per case from arch taper and arch closure instead of the stated RAS convention, which does not hold across the release, and each arch is reduced to an occlusal ridge profile in arch-angle coordinates yielding overbite, overjet, midline deviation, transverse overlap, crossbite extent, cusp interdigitation lag, and the occlusal curves in closed form. Gradient boosting maps 31 such measurements onto 13 template fields, a field being predicted only where patient-level cross-validation beats its own majority baseline, and a deterministic renderer emits the corpus six-part narrative; a ConvNeXt-Tiny classifier over the five standardised photographic views is fused per field, raising mean field accuracy from 0.601 to 0.683. Reimplementation of the challenge evaluator shows that its BLEU-4 and METEOR are local variants whose F-mean weights recall nine to one, that two clinicians agree on 47 percent of findings for the same patient, and that a constant report consequently outscores a genuine second clinician report by 0.165 captioning. Held-out scores reach BLEU-4 0.458 and METEOR 0.677 against intraoral scan references and 0.278 and 0.507 against photograph references, and the submitted system placed third in the ODIN 2026 Bite2Text test phase at 0.2680 and 0.4629, within 0.022 BLEU-4 of first, running on CPU in under ten seconds per case. The dataset and code are available at https://github.com/GIND123/ODIN_toothfairy4
Comments10 pages, 4 figures. Third-place system in the ODIN 2026 Bite2Text test phase. Code and data processing resources: https://github.com/GIND123/ODIN_toothfairy4