arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05728cs.CL

MedWER:一种可复现、无模型的医学语音识别评估协议

MedWER: A Reproducible, Model-Free Evaluation Protocol for Medical Speech Recognition

发表机构诺迪斯科技有限公司
查看机构详情
  • NORDIS TECH INC.(诺迪斯科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Justin Behling

首次发表
浏览论文内容

中文总结 AI 辅助

MedWER提出一种可复现、无模型的医学语音识别评估协议,通过固定术语列表和规范化器,解决词错误率掩盖临床错误的问题,并在两个开放基准上验证了其有效性。

中文摘要 AI 辅助

整体词错误率掩盖了临床关键错误:一段转录文本可能达到95%的正确率,却仍然将一种药物替换为另一种。通常的解决方案是对医学实体上的错误进行加权,且几乎总是依赖于评估时的命名实体识别(NER)模型或云API,这使得指标的分母成为一个版本化的黑盒。我们提出了MedWER,一种用于医学自动语音识别(ASR)的评估协议和开源工具,其分母是一个固定的、许可干净的术语列表:从公共来源投影得到的19,373个药物、诊断、症状和损伤机制条目。该协议将一个固定的文本规范化器与一个短语感知的术语受限WER(即MedWER)相结合,因此唯一版本化的组件是一个保持在精确版本并对照已提交的黄金测试夹具进行校验的规范化器依赖。覆盖率针对一个独立的省级药物福利文件进行验证,该文件并非用于构建列表;匹配启发式方法则针对真实实体跨度进行校准。使用发布的工具对Moonshine base、Whisper tiny.en和MedASR在两个开放基准上的基线进行评分,并通过重采样的逐话语得分报告95%置信区间。

英文摘要

Overall word error rate hides clinically critical errors: a transcript can be 95% correct and still swap one drug for another. The usual fix weights errors on medical entities, and almost always depends on an evaluation-time named-entity recognition (NER) model or cloud API, which makes the metric's denominator a versioned black box. We present MedWER, an evaluation protocol and open-source tool for medical ASR whose denominator is a fixed, license-clean term list: 19,373 drug, diagnosis, symptom, and injury-mechanism entries projected from public sources. The protocol couples a pinned text normalizer with a phrase-aware term-restricted WER, the MedWER, so the only versioned component is a normalizer dependency held at an exact release and checked against committed golden fixtures. Coverage is validated against an independent provincial drug-benefit file the list was not built from; the matching heuristic is calibrated against ground-truth entity spans. Baselines for Moonshine~base, Whisper~base.en, and MedASR on two open benchmarks are scored with the released tool and reported with 95% confidence intervals from resampled per-utterance scores.

补充信息

↑