KnowBench:以工作量减少作为临床AI的统一、部署落地基准
KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI
- Knowtex Inc.(Knowtex 公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
KnowBench以工作量减少为核心指标,统一评估临床AI在真实部署中的效率,实测总体ER达97.99%,为跨系统比较提供可审计标准。
AI中文摘要:
临床AI系统目前使用的评估工具是为研究场景而构建的(基于参考的相似度指标和专家评分面板),这些工具衡量的是与人工产物的相似程度,而非实际负担的减轻。我们推出由Knowtex首创的KnowBench,其统一指标为工作量减少(ER):即经过负责任临床医生在专家和安全审查下接受的系统生成的临床工作产品的比例。ER被统一定义,并在临床AI自动化的行政工作负载中按任务实例化:就诊记录、诊断和计费编码、医嘱、电子健康记录(EHR)图表总结、患者诊后总结以及临床决策支持。在每一个实例化中,构建方式完全相同:临床医生的审查和认证事件是金标准,每个被接受的工作单元都是系统完成的工作,每次修正都是返还给临床医生的剩余工作量。本文的主要贡献在于基准本身:指标、其退化情形,以及一套报告协议,在该协议下ER声明可被审计且跨系统可比。此外,我们报告了来自文档实例化的初步首要测量结果:在超过六个月的生产窗口和十三个医学专科中,超过一百万个已签署的就诊记录中,Knowtex专有的微调临床基础模型在闭环反馈架构内运行,实现了97.99%的总体ER,各专科的汇总范围在96.8%至98.9%之间。本发布部分报告了协议检查清单,并说明哪些配套统计数据被保留;提供该基准是为了让这一数字以及之后报告的每一个数字都能遵循相同的标准。
英文摘要:
Clinical AI systems are evaluated with instruments built for research settings (reference-based similarity metrics and expert rubric panels) that measure resemblance to an artifact rather than reduction of a burden. We introduce KnowBench, pioneered by Knowtex, whose unifying metric is Effort Reduction (ER): the proportion of system-generated clinical work product accepted by the responsible clinician under expert and safety review. ER is defined once and instantiated per task across the administrative workload clinical AI automates: visit notes, diagnosis and billing codes, orders, EHR chart summarization, patient after-visit summaries, and clinical decision support. In every instantiation the construction is identical: the clinician's review-and-attestation event is the ground truth, every accepted unit is work the system completed, and every correction is residual effort returned to the clinician. The primary contribution of this paper is the benchmark itself: the metric, its degenerate cases, and a reporting protocol under which ER claims are auditable and cross-system comparable. Alongside it we report an initial headline measurement from the documentation instantiation: over one million signed encounters across a production window exceeding six months and thirteen medical specialties, Knowtex's proprietary fine-tuned clinical foundation models operating inside a closed feedback architecture achieve an aggregate ER of 97.99%, with per-specialty aggregates spanning 96.8-98.9%. This release reports the protocol's checklist partially, and states which companion statistics are withheld; the benchmark is offered so that this figure, and every figure reported after it, can be held to the same standard.