CoDA:探索链式分布攻击及事后令牌空间修复用于医疗视觉-语言模型
CoDA: Exploring Chain-of-Distribution Attacks and Post-Hoc Token-Space Repair for Medical Vision-Language Models
浏览论文内容
中文总结 AI 辅助
本文提出CoDA框架,通过构建临床合理的流程偏移,评估医疗视觉-语言模型在真实临床工作流中的可靠性,发现零样本性能显著下降,并引入事后修复策略提升准确性。
中文摘要 AI 辅助
医疗视觉-语言模型(MVLMs)在放射学流程中被广泛用作感知骨干,并作为多模态助手的视觉前端,但其在真实临床工作流中的可靠性仍待探索。以往的鲁棒性评估常假设干净、经过整理的输入或研究孤立的损坏,忽略了常规的获取、重建、显示和交付操作,这些操作在保持临床可读性的同时改变图像统计特性。为填补这一空白,我们提出了CoDA,一种链式分布框架,通过组合类似获取的阴影、重建和显示的重映射,以及交付和导出的损坏来构建临床合理的流程偏移。在受掩蔽的结构相似性约束下,CoDA联合优化阶段组成和参数以诱导失败,同时保持视觉合理性。在脑部MRI、胸部X光和腹部CT中,CoDA显著降低了CLIP风格MVLMs的零样本性能,链式组合比任何单一阶段更具破坏性。我们还评估了多模态大语言模型(MLLMs)作为影像真实性和质量的技术-真实性审计员,而非病理学。专有多模态模型在CoDA偏移样本上表现出降级的审计可靠性和持续的高置信度错误,而我们测试的医疗专用MLLMs在医学图像质量审计中表现出明显缺陷。最后,我们引入了一种基于教师引导的令牌空间适应的后置修复策略,通过补丁级对齐提高已存档CoDA输出的准确性。总体而言,我们的发现表征了MVLM部署的临床基础威胁面,并展示了轻量级对齐提高了部署的鲁棒性。
英文摘要
Medical vision--language models (MVLMs) are increasingly used as perceptual backbones in radiology pipelines and as the visual front end of multimodal assistants, yet their reliability under real clinical workflows remains underexplored. Prior robustness evaluations often assume clean, curated inputs or study isolated corruptions, overlooking routine acquisition, reconstruction, display, and delivery operations that preserve clinical readability while shifting image statistics. To address this gap, we propose CoDA, a chain-of-distribution framework that constructs clinically plausible pipeline shifts by composing acquisition-like shading, reconstruction and display remapping, and delivery and export degradations. Under masked structural-similarity constraints, CoDA jointly optimizes stage compositions and parameters to induce failures while preserving visual plausibility. Across brain MRI, chest X-ray, and abdominal CT, CoDA substantially degrades the zero-shot performance of CLIP-style MVLMs, with chained compositions consistently more damaging than any single stage. We also evaluate multimodal large language models (MLLMs) as technical-authenticity auditors of imaging realism and quality rather than pathology. Proprietary multimodal models show degraded auditing reliability and persistent high-confidence errors on CoDA-shifted samples, while the medical-specific MLLMs we test exhibit clear deficiencies in medical image quality auditing. Finally, we introduce a post-hoc repair strategy based on teacher-guided token-space adaptation with patch-level alignment, which improves accuracy on archived CoDA outputs. Overall, our findings characterize a clinically grounded threat surface for MVLM deployment and show that lightweight alignment improves robustness in deployment.