发表机构
British Columbia Cancer Registry, Provincial Health Services Authority; University of British Columbia; Data Science Institute, University of British Columbia(不列颠哥伦比亚癌症登记处,省级卫生服务管理局; 英属哥伦比亚大学; 英属哥伦比亚大学数据科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对癌症登记中报告级人工标注训练数据稀缺问题,利用操作生成的患者级标签,通过基于注意力的多实例学习恢复标签与报告的联系,开发框架训练深度学习分类器,提升肿瘤组分类性能。
AI 中文摘要
利用深度学习使癌症登记现代化为自动化病理报告编码带来新机遇,但受报告级人工标注训练数据稀缺限制。癌症登记产生大量患者级专家分配标签,却与相关病理报告无关联。我们开发框架,用基于注意力的多实例学习恢复联系,训练分类器性能良好,为自动化癌症登记工作流程提供实用途径。
英文摘要
Modernizing cancer registries with deep learning is opening new opportunities to automate labor-intensive tasks such as the coding of pathology reports. However, progress is constrained by the scarcity of report-level human-annotated training data. Cancer registries generate substantial volumes of expert-assigned labels as a routine product of their operations, but these exist at the patient level and are not linked to the individual pathology reports that informed them, limiting their direct use for training models. We develop an efficient framework for training deep learning classifiers by leveraging these operationally-generated labels without requiring per-report human annotation, demonstrated for tumor group classification at the BC Cancer Registry. We use Attention-Based Multiple Instance Learning (ABMIL) to recover the lost link between patient-level labels and the reports that informed them, leveraging the attention the model places on each report to distil a large, noisily-labeled corpus into a compact, high-quality per-report training dataset. A classifier fine-tuned on a distilled dataset achieved a macro F1 of 0.83, outperforming established baselines across most tumor groups. By turning routine operational labels into high-quality training data without additional annotation or large-scale computing infrastructure, ABMIL offers a practical and accessible route to automating cancer registry workflows.