woma:用于内窥镜的实时基础模型及其微调模型
woma: a real-time foundation model and its fine-tuned models for endoscopy
- CloudKites AI Lab(云筝人工智能实验室)
- Monash Business School, Monash University(莫纳什大学莫纳什商学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
woma 是一种无需标签训练的内窥镜实时基础模型,通过系统化生产设计微调出结肠镜和胃镜模型,在多项基准上达到高精度,并以约100帧/秒的速度运行,优于主流推理框架。
AI中文摘要:
woma 是一个用于胃肠道内窥镜的实时基础模型:一个在约一百万帧内窥镜图像上无需标签训练的网络,任务模型从中进行微调。我们为生产环境贡献了一种系统化设计。在任何运行之前就确定了要求和通过标准,根据预注册规则筛选了八个候选方案,自监督训练遵循停止规则,随后进行微调和部署优化,所有工作都在一个自包含的库 numbat 中完成。我们还贡献了 woma 本身及其两个微调模型,报告了每个结果是否达标。我们的结肠镜模型能够发现并勾勒息肉,识别视野中的结肠段,建议息肉类型并评估肠道准备质量。我们的胃镜模型能够从22个协议站点中命名一个站点,标记并勾勒病变,并命名七种发现之一。每个数字都是在训练中从未见过的数据上读取的,发布的权重基于该记录选择。在结肠镜检查中,在包含六家医院的 PolypGen 数据集中,96% 的息肉以精度≥0.85 被发现,在十五个完整的 REAL-Colon 视频中,19 个息肉全部被发现,每次操作仅出现1.6次误报。在胃镜检查中,92% 的未见患者帧中地标区域被正确命名,39 个保留的肿瘤帧中有37个被标记,特异性为0.91。在一台工作站 GPU 上,每个任务在1080p视频上以约每秒100帧的速度运行,在所有四种测试的精度设置下均快于 PyTorch、ONNX Runtime 和 TensorRT。TensorRT 最为接近:其运行一次我们的基础模型所需时间比我们长3%至27%,而从帧到结果,我们每秒多提供6%至31%的帧。另一个构建完全不链接任何供应商库——我们自己的基于 Vulkan 的内核——因此站点部署两个文件即可,无需工具包、无需 cuDNN、无需框架;在 f32 精度下,它在同一张显卡上优于 CUDA 构建。
英文摘要:
woma is a real-time foundation model for gastrointestinal endoscopy: a network trained without labels on about a million endoscopy frames, from which task models are fine-tuned. We contribute a systematic design for production. Requirements and pass marks were fixed before any run, eight candidates screened under pre-registered rules, self-supervised training taken to a stopping rule, then fine-tuning and deployment optimisation, all on one self-contained library, numbat. We also contribute woma itself with two fine-tuned models, every outcome reported met or missed. Our colonoscopy model finds and outlines polyps, names which colon segment is in view, suggests polyp type and grades bowel preparation. Our gastroscopy model names a station out of 22 protocol sites, flags and outlines lesions, and names one of seven findings. Every number was read on data never seen in training, and shipped weights were chosen on that record. In colonoscopy, 96% of polyps in a six-hospital PolypGen set are found at precision >=0.85, and 19 of 19 polyps across fifteen full REAL-Colon videos at 1.6 false alarms per procedure. In gastroscopy, landmark region is named correctly on 92% of frames from unseen patients, and 37 of 39 held-out neoplasia frames are flagged at specificity 0.91. On one workstation GPU every task runs over 1080p video at about 100 frames per second, faster than PyTorch, ONNX Runtime and TensorRT in all four precision regimes tested. TensorRT comes closest: one pass of our foundation model takes it 3 to 27% longer than ours, and we deliver 6 to 31% more frames per second from frame to results. A second build links no vendor library at all -- our own kernels over Vulkan -- so a site deploys two files and needs no toolkit, no cuDNN and no framework; in f32 it beats the CUDA build on the same card.