发表机构
The Hong Kong Polytechnic University; Beihang University; Hangzhou International Innovation Institute, Beihang University; Nanjing University of Aeronautics and Astronautics; Shandong University(香港理工大学; 北京航空航天大学; 北京航空航天大学杭州国际创新研究院; 南京航空航天大学; 山东大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出受控复杂度审计框架,对比Full-EDL与Simple-CE+TS等模型,发现Simple-CE+TS在两数据集表现更优,选择其作为分布内目标,Full-EDL作参考,强调组件保留需功能与重训练证据。
AI 中文摘要
多器官超声分类器正越来越多地结合注意力机制、混合专家路由、不确定性门控以及证据深度学习(EDL)目标,以应对异质解剖结构和采集差异。然而,看似合理的设计理由本身并不能证明新增组件能提升训练后的系统性能。本文提出一种受控复杂度审计框架,将其应用于最大证据候选模型Full-EDL与更简单替代方案之间的部署决策。在主数据集上评估了六个候选模型,在内部复现研究中评估了三个候选模型,采用十组匹配随机种子、冻结图像级划分、考虑容量与优化的比较、对称温度缩放、配对决策规则以及单独的分布外(OOD)弃权(不执行)机制。保留Full-EDL在两个数据集上均未获得可靠的宏F1值提升,而简化替代方案在非劣效性边际下仍无定论。带温度缩放的简单交叉熵(Simple-CE+TS)在两个数据集上均满足校准负对数似然准则,并表现出良好的选择性风险排序。证据训练的原始校准优势在温度缩放后消失,且未在第二个数据集上重现。门控机制在审计检查点处的可观测影响可忽略不计,删除仅Full的链未显示出稳定的任务或校准损失收益。Simple-CE触发了针对胎儿探头的OOD弃权(不执行),但未触发针对肺部探头的,因此无法做出无条件的OOD安全性声明。因此,本文选择Simple-CE+TS作为评估的分布内目标,同时保留Full-EDL作为最大参考。组件应通过基于功能和重新训练的证据获得保留资格,且校准与分布偏移可靠性应分别进行评估。
英文摘要
Multi-organ ultrasound classifiers increasingly combine attention, mixture-of-experts routing, uncertainty gating, and evidential deep learning (EDL) objectives to address heterogeneous anatomy and acquisition. Yet a plausible design rationale does not by itself establish that an added component improves the trained system. We contribute a controlled complexity-audit framework, applied to the deployment decision between the maximal evidential candidate Full-EDL and simpler alternatives. Six candidates were evaluated on the primary dataset and three in an internal replication, using ten matched seeds, frozen image-level partitions, capacity- and optimisation-aware comparisons, symmetric temperature scaling, paired decision rules, and a separate out-of-distribution (OOD) veto. Retaining Full-EDL did not establish a reliable macro-F1 gain on either dataset, while the simplified alternatives remained inconclusive under the non-inferiority margin. Simple cross-entropy with temperature scaling (Simple-CE+TS) met the calibrated negative log-likelihood criterion on both datasets and showed favourable selective-risk ordering. The raw calibration advantage of evidential training disappeared after temperature scaling and did not recur on the second dataset. The gate had negligible observable influence at the audited checkpoints, and deleting the Full-only chain revealed no stable task or calibrated-loss benefit. Simple-CE nevertheless triggered the OOD veto against the fetal probe but not the lung probe, precluding an unconditional OOD-safety claim. We therefore selected Simple-CE+TS for the evaluated in-distribution objective while retaining Full-EDL as the maximal reference. Components should earn retention through functional and retraining-based evidence, and calibration and distribution-shift reliability should be evaluated separately.
CommentsSupplementary material is available as an ancillary file on this arXiv page