发表机构
ETH Zurich; Berlin Institute of Health at Charité – Universitätsmedizin Berlin; Deutsches Herzzentrum der Charité – Medical Heart Center of Charité and German Heart Institute Berlin(苏黎世联邦理工学院; 柏林夏里特医学院柏林健康研究所; 柏林夏里特医学院德国心脏中心与柏林德国心脏研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究LLM水印对医学性能的影响,在11个LLM和7个VLM上对5种水印方案进行多任务基准测试,引入人类专家验证管道补充评估,发现水印会致多种退化,特定领域评估是医学中水印模型安全部署的前提。
AI 中文摘要
大语言模型(LLMs)越来越多地集成到临床工作流程中,因此需要通过水印对模型生成的输出进行可靠溯源。然而,大多数水印是在通用基准上评估的,医学等领域未得到充分探索,在这些领域小的词元级扰动可能导致重大语义变化。本文首次对LLM水印如何影响医学性能进行了严谨研究,在11个LLM和7个VLM上对5种水印方案进行跨单峰和多峰临床推理任务的基准测试。通过引入经人类专家验证的管道系统地审核医学推理质量、术语准确性和幻觉,补充现有评估。结果表明水印会导致多种故障模式的显著退化,包括词汇损坏、幻觉术语以及图像发现的错误归因或遗漏加剧。缺乏特定领域分析和忽略临床文本固有故障的综合指标会掩盖水印导致的实际退化。研究结果表明特定领域评估是医学中水印模型安全部署的前提,当前基准可能掩盖临床后果严重的故障。
英文摘要
Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-generated output with watermarking. Yet, most watermarks are evaluated on general-purpose benchmarks, leaving domains like medicine, where small token-level perturbations can result in significant semantic changes, underexplored. In this work, we present the first rigorous study of how LLM watermarks affect medical performance, benchmarking 5 watermarking schemes across 11 LLMs and 7 VLMs on various tasks spanning unimodal and multimodal clinical reasoning. Importantly, we complement existing evaluations by introducing a human-expert-validated pipeline for systematically auditing medical reasoning quality, terminological precision, and induced hallucinations. Our results reveal that watermarking can induce substantial degradation across multiple failure modes, including lexical corruption, hallucinated terminology, and amplified misattribution or omission of image findings. Notably, we find that the absence of domain-specific analyses, combined with aggregate metrics that miss failures inherent to clinical text, can systematically obscure practical watermark-induced degradations. Our findings establish domain-specific evaluation as a prerequisite for the safe deployment of watermarked models in medicine, where current benchmarks can otherwise mask clinically consequential failures.