MIRROR:多模态智能放射学推理与观察报告生成器
MIRROR: Multimodal Intelligent Radiology Reasoning and Observation Reporter
- Texas A&M University(德克萨斯农工大学)
- Capital One(第一资本金融公司)
- IIT Hyderabad(印度理工学院海得拉巴分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MIRROR是一款多模态智能放射学报告生成原型,通过分离分类、定位与文本生成环节实现可审计的放射学发现,在ChestMNIST数据集上达到0.729的宏平均AUROC,但受类别不平衡影响存在决策弃权问题。
AI中文摘要:
放射科医生读取模型输出时面临两个问题:模型仅返回数值却不给出理由,任何将该数值转换为可读文本的系统都可能悄悄加入模型从未做出的断言。MIRROR是一款用于区分这些缺陷的研究原型,它串联了多标签分类器、将每个阳性发现转换为命名解剖区域的Grad-CAM定位器,以及接收标签、概率和区域但不接收图像的报告生成器。由于语言层无法看到像素,它无法断言分类器未做出的发现。我们明确说明了这一设计的优势:MIRROR报告的发现可针对概率向量进行审计,而构建发现的句子是普通生成文本,我们展示了一份报告提及系统从未测量的心胸比率。一个注册表存储胸部X射线、脑部MRI和头部CT的分类学、解剖学和措辞,因此添加模态只需进行数据变更;三个模态均被路由并测试,其中一个接受训练。在ChestMNIST数据集上,该分类器的宏平均AUROC达到0.729,且在全部14个标签上的表现均优于随机水平,其精度是随机排序器的1.6至6.8倍。然而在默认0.5阈值下,它对其中11个标签完全未输出阳性预测,其出色的Brier评分为0.045,却与忽略图像的预测器获得的0.047相当。这种区分能力是真实的,但决策并非如此。在放射学常见的类别不平衡情况下,聚合指标会高估未采取任何行动的模型,应将其与该基准进行对比报告。
英文摘要:
A radiologist reading a model's output faces two problems. The model returns a number and no reason, and any system that turns that number into readable prose can quietly add claims the model never made. MIRROR is a research prototype built to separate those failures. It chains a multi-label classifier, a Grad-CAM localizer that turns each positive finding into a named anatomical region, and a report writer that receives the labels, probabilities, and regions but never the image. Because the language layer cannot see pixels, it cannot assert a finding the classifier did not make. We are precise about what that buys: a MIRROR report's findings are auditable against the probability vector, while the sentences framing them are ordinary generated text, and we show one stating a cardiothoracic ratio the system never measured. One registry holds the taxonomy, anatomy, and phrasing for chest X-ray, brain MRI, and head CT, so adding a modality is a data change; all three are routed and tested, one is trained. On ChestMNIST that classifier reaches macro AUROC 0.729 and ranks better than chance on all 14 labels, at 1.6 to 6.8 times the precision a random ranker would get. Yet at the default 0.5 threshold it emits no positive prediction at all for 11 of them, and its excellent-looking Brier score of 0.045 sits beside the 0.047 earned by a predictor that ignores the image. The discrimination is real; the decisions are not. Under the class imbalance normal in radiology, aggregate metrics flatter models that do nothing, and should be reported against that floor.