arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26886cs.CVcs.AIcs.CLcs.CY

Hearsay:无需图像的视觉-语言医学诊断

Hearsay: Vision-Language Medical Diagnoses Without an Image

Siddharth Vohra

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现前沿视觉-语言模型在无图像时会基于人口统计学特征编造医学诊断,存在不同失败模式,提出需直接审计其结构化输出通道并将探针词敏感性作为核心评估维度。

中文摘要 AI 辅助

当被要求描述一张从未提供的医学图像时,前沿视觉-语言模型(VLM)不会弃权(不执行),而是会编造诊断结果。我们发现这种编造并非随机,而是由所述患者的特征决定的。在胸部X光、脑部MRI和皮肤病学领域,仅向Claude Opus-4.7、GPT-5.4和Gemini-3.1-Pro提供人口统计学描述符而不提供图像,改变描述符会系统性地改变返回的诊断结果。Claude的表现高度集中:一名65岁白人男性询问皮肤痣时,几乎所有回复都给出黑色素瘤诊断;一名32岁黑人女性询问胸部X光时,得到结节病诊断,其推理为“基于人口统计学特征和典型模式疑似”。GPT-5.4的影响更广泛,在我们测试的所有人口统计学类别中均会编造诊断,最显著的是为胸部X光检查的年轻黑人患者命名结节病。两项结构性发现加剧了该问题:存在一种模糊机制,即文本承认图像缺失,但结构化诊断字段仍会命名疾病,这种分离在仅文本的审计中不可见;当“皮肤痣”替换为“皮肤病变”时,Claude的皮肤病学效应完全消失,而GPT-5.4的效应仍存在,表明这种幻象是一系列不同的失败模式而非单一现象。临床流程中可信赖的VLM部署需要直接审计结构化输出通道,且探针词敏感性应被视为一级评估维度。

英文摘要

When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts the diagnosis returned. Claude concentrates sharply: a 65-year-old white man asking about a skin mole receives Melanoma in nearly every response, and a 32-year-old Black woman asking about her chest X-ray receives a Sarcoidosis diagnosis whose reasoning reads "suspected, based on demographics and classic pattern.'' GPT-5.4's effect is broader, fabricating across every demographic cell we test, most conspicuously naming Sarcoidosis for young Black patients on chest X-ray. Two structural findings sharpen the problem. A hedged regime appears in which the prose acknowledges the missing image while the structured diagnosis field nevertheless names a disease, a dissociation invisible to prose-only audits. And Claude's dermatology effect collapses entirely when 'skin mole' is swapped for 'skin lesion' while GPT-5.4's is preserved, indicating that mirage is a family of distinct failure modes rather than a single phenomenon. Trustworthy VLM deployment in clinical pipelines requires auditing the structured output channel directly, and probe-word sensitivity should be treated as a first-class evaluation dimension

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • Amazon Web Services AI Native(亚马逊网络服务AI原生部门)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑