arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25589cs.CVcs.AIcs.CL

放射学视觉语言模型基准的法证可重复性审计:从预期协议到发布工件

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

Mateusz Kozłowski

首次发表
浏览论文内容

中文总结 AI 辅助

对胸部X光视觉语言模型试点进行法证可重复性审计,追踪提示绑定等多方面情况,发现存在图像渲染、数据分割等问题,重建队列改变统计值,撤回原声明并指定机器可验证控制。

中文摘要 AI 辅助

医学成像人工智能基准结合了数据集、DICOM 渲染、提示、提供者应用程序编程接口、自动标签、统计代码、手稿和存储库发布。通常假定这些工件之间是一致的,而不是进行测试。我们对一个保存的胸部 X 光视觉语言模型(VLM)试点进行了回顾性法证可重复性审计;没有再次调用模型,也没有新注释图像或报告。我们追踪了提示绑定、DICOM 元数据、输出完整性、标签提取、匹配分析和发布传播。在 300 个计划的模型 - 提示调用中,297 个产生了非空报告。使用相同的 C 提示执行了 60 次标记为 A/B 的克劳德调用。30 项研究代表 28 名患者。4 张 MONOCHROME1 图像在没有进行所需极性反转的情况下进行了渲染,数据集分割成员未保留,未经验证的提取器将 5 份报告截断为 4000 个字符。重建一个由 369 个完整病例发现块组成的共同队列,将 Cochr an 的 Q 从 154.73 变为 182.29。在 45 次 McNemar 比较中,27 次未调整的 p < 0.05,20 次在 Holm 调整后仍低于 0.05。这些值仅描述存档的自动标签矩阵;它们无法恢复预期的提示比较或确定临床性能。我们撤回原始的性能、排名、提示效果和临床声明,并为队列、DICOM 渲染、提示和模型标识、调用状态、注释来源、键控分析和派生工件指定机器可验证的控制。

英文摘要

Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases. Agreement across these artifacts is usually assumed rather than tested. We performed a retrospective forensic reproducibility audit of a preserved chest-radiograph vision-language model (VLM) pilot; no model was called again and no image or report was newly annotated. We traced prompt bindings, DICOM metadata, output completeness, label extraction, matched analyses, and release propagation. Of 300 planned model-prompt calls, 297 yielded nonempty reports. Sixty Claude calls labeled A/B were executed with the same C prompt. The 30 studies represented 28 patients. Four MONOCHROME1 images were rendered without required polarity inversion, dataset split membership was not retained, and the unvalidated extractor truncated five reports to 4000 characters. Reconstructing one common cohort of 369 complete case-finding blocks changed Cochran's Q from 154.73 to 182.29. Of 45 McNemar comparisons, 27 had unadjusted p < 0.05 and 20 remained below 0.05 after Holm adjustment. These values describe only the archived automated-label matrix; they do not recover the intended prompt comparison or establish clinical performance. We withdraw the original performance, ranking, prompt-effect, and clinical claims and specify machine-verifiable controls for cohort, DICOM rendering, prompt and model identity, call status, annotation provenance, keyed analysis, and derived artifacts.

补充信息

↑