衡量VLMs不表达的内容:验证指标隐藏了放射学报告生成中的临床术语擦除
Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
- DTU Compute, Technical University of Denmark(丹麦技术大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出临床关联位移(CAD)和加权关联擦除(WAE)指标,用于评估放射学报告生成中临床术语的保留情况,揭示当前验证指标在临床术语擦除方面的不足。
AI中文摘要:
可靠地将视觉语言模型(VLMs)应用于放射学需要超越表面文本相似度的验证指标,以确保临床准确性和人口统计学公平性。本文研究了当前模型评估中的一个关键盲点:解码策略的使用导致高聚合词重叠分数,尽管这些策略导致模板崩溃,即模型只生成重复的、安全的通用文本并省略临床术语。未解决这一盲点可能导致指标游戏,即在基准测试中表现良好的模型在临床上信息不足。相反,我们倡导使用词汇多样性度量来检查模型生成的临床特异性。我们引入了临床关联位移(CAD),一种词汇层面的框架,用于量化生成报告中基于人口统计学的词关联变化。加权关联擦除(WAE)将这些变化汇总,以衡量跨人口统计学群体的临床信号损失。我们证明确定性解码会产生高水平的语义擦除,而随机采样生成多样化的输出但可能引入新的偏见,这促使我们重新思考“最优”报告的定义。
英文摘要:
Reliable deployment of Vision-Language Models (VLMs) in radiology requires validation metrics that go beyond surface-level text similarity to ensure clinical fidelity and demographic fairness. This paper investigates a critical blind spot in current model evaluation: the use of decoding strategies that lead to high aggregate token-overlap scores despite succumbing to template collapse, in which models generate only repetitive, safe generic text and omit clinical terminology. Unaddressed, this blind spot can lead to metric gaming, where models that perform well on benchmarks prove clinically uninformative. Instead, we advocate for lexical diversity measures to check model generations for clinical specificity. We introduce Clinical Association Displacement (CAD), a vocabulary-level framework that quantifies shifts in demographic-based word associations in generated reports. Weighted Association Erasure (WAE) aggregates these shifts to measure the clinical signal loss across demographic groups. We show that deterministic decoding produces high levels of semantic erasure, while stochastic sampling generates diverse outputs but risks introducing new bias, motivating a fundamental rethink of how "optimal" reporting is defined.