arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15015cs.AI

四本账本,而非一个分数:生物医学机器学习中LLM评判校准的责任沟通

Four Ledgers, Not One Score: Responsible Communication of LLM-Judge Calibration in Biomedical ML

Sidi Chang, Peiying Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过审计一个私有工作流,指出LLM评判器校准中混淆不同账本的问题,提出来源感知审计、存储契约和最小校准门,以负责任地沟通生物医学ML能力主张。

中文摘要 AI 辅助

合成扰动似乎为生物医学机器学习中的LLM评估器提供了廉价的校准数据,而该领域专家评审稀缺。然而,植入的突变密钥既不是检测器输出,也不自动等同于人类真实标签。我们形式化了四个不同的账本:植入扰动、独立检测器输出、来源关联的人类处置以及人类新增发现。随后,我们审计了一个私有合成日语护理交接工作流的评估设计、评分代码、读取路径和当前人类记录。该工厂在47个目标中存储了69张植入错误卡片。最终评审覆盖22个目标,包含22个已确认的导入提案、9个被拒绝的提案和79张人类新增卡片;仅有3个评审目标进行了双重标注。将导入的植入密钥传递给通用检测器评分器,得到22/(22+9)=0.710和22/(22+79)=0.218。直接审计恒等式表明,这些值是提案确认产出和提交账本构成,而非评判器的精确率和召回率,因为在可用记录中,对于被审计的提案,没有保留独立的检测器实现。审计还发现来源名称冲突、行遮蔽、强制严重性、空洞比率默认值以及不支持的零支持字段权重。我们贡献了来源感知的主张审计、存储契约和最小校准门,以负责任地沟通生物医学机器学习能力主张。这一单工作流取证案例是失败模式的存在性证明,而非其普遍性的估计:现有的人类工作支持对合成提案的探索性审计,但不支持LLM评判器操作特性、临床有效性、语料库普遍性或稳健的注释者间一致性。

英文摘要

Synthetic perturbations appear to offer inexpensive calibration data for LLM evaluators in biomedical ML, where expert review is scarce. Yet a planted mutation key is neither a detector output nor automatically human ground truth. We formalize four distinct ledgers: planted perturbations, independent detector outputs, source-linked human dispositions, and human-added discoveries. We then audit the evaluation design, scoring code, read paths, and current human records of a private synthetic Japanese care-handoff workflow. The factory stored 69 planted error cards across 47 targets. Final review covers 22 targets and contains 22 confirmed imported proposals, 9 rejected proposals, and 79 human-added cards; only 3 reviewed targets are double annotated. Passing imported plant keys to a generic detector scorer yields 22/(22+9)=0.710 and 22/(22+79)=0.218. A direct audit identity shows that these values are proposal-confirmation yield and submitted-ledger composition, not judge precision and recall, because no independent detector realization was preserved for the audited proposals in the available records. The audit also finds source-name collisions, row shadowing, forced severity, vacuous ratio defaults, and unsupported zero-support field weights. We contribute a provenance-aware claim audit, a storage contract, and a minimum calibration gate for responsibly communicating biomedical ML capability claims. This single-workflow forensic case is an existence proof of a failure mode, not an estimate of its prevalence: existing human work supports an exploratory audit of synthetic proposals, but not LLM-judge operating characteristics, clinical validity, corpus prevalence, or robust inter-annotator agreement.

发表机构

  • Blossom AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑