AI 中文总结
CARE-X是一款胸部X射线视觉语言模型,通过辅助监督、奖励对齐学习及工具增强测量,在多个医学影像任务上实现了优于基线的性能,缩小了放射科医生需求与现有生成模型间的差距。
AI 中文摘要
临床可用的胸部X射线系统必须超越流畅的报告生成:它应能以可调决策阈值对影像发现进行分类、在空间上定位这些发现,并得出许多诊断所依赖的解剖测量值。如今的视觉语言模型(VLMs)即便处理这些任务,也会将其视为独立问题,导致放射科医生的需求与生成模型提供的功能之间存在差距。我们推出了CARE-X,这是一款胸部X射线VLMs,通过将辅助判别监督与奖励对齐生成相结合来缩小这一差距。CARE-X在其生成主干上增加了焦点损失分类和复合损失定位头,与语言建模目标共同训练。这种辅助监督产生了带有可调决策阈值的判别式诊断预测和精确的空间定位,同时还提升了报告质量,提供了结构化预测与生成相互增强的证据。在此基础上,解耦剪辑与动态采样策略优化(DAPO)利用针对报告生成、视觉问答(VQA)和空间定位的特定任务奖励信号,直接优化实践中重要的临床质量指标。其结果是在四个报告生成基准的大多数指标上达到了最先进的性能,在ReXVQA上的VQA准确率为94.0%(比次优基线高出6.0个百分点),生成式空间解码达到了与专用检测头近乎相当的水平。此外,为解决依赖测量的诊断问题,我们将Qwen3-VL-4B-Instruct与原生工具调用能力相结合,用于调用确定性测量工具,同时保留对图像的完整视觉访问权限。这种混合推理在五种依赖测量的条件下,比仅感知基线的平均F1值高出43.6个百分点。
英文摘要
A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend. Today's Vision-Language Models (VLMs) treat these as separate problems, if they address them at all, leaving a gap between what radiologists need and what generative models provide. We introduce CARE-X, a chest X-ray VLM that narrows this gap by unifying auxiliary discriminative supervision with reward-aligned generation. CARE-X augments its generative backbone with focal-loss classification and composite-loss grounding heads, co-trained alongside the language-modeling objective. This auxiliary supervision produces discriminative diagnostic predictions with tunable decision thresholds and precise spatial localization while also improving report quality, providing evidence that structured prediction and generation reinforce one another. Building on this foundation, Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) leverages task-specific reward signals for report generation, visual question answering (VQA), and spatial grounding, directly optimizing the clinical quality metrics that matter in practice. The result is state-of-the-art performance on the majority of metrics across four report-generation benchmarks, 94.0% VQA accuracy on ReXVQA (+6.0 pp over the next-best baseline), and generative spatial decoding that reaches near parity with dedicated detection heads. Separately, to address measurement-dependent diagnoses, we couple Qwen3-VL-4B-Instruct with native tool-calling capabilities for invoking deterministic measurement tools, while retaining full visual access to the image. This hybrid inference yields +43.6 pp average F1 over perception-only baselines across five measurement-dependent conditions.