arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

做($x$)还是不做($x$):用于数据集增强的医学图像反事实

To do($x$) or not to do($x$): Medical Image Counterfactuals for Dataset Augmentation

Yasin Ibrahim, Robin J. Evans, Konstantinos Kamnitsas

arXiv 2609.14124首次发表:更新:

发表机构

University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究区分医学图像反事实生成的因果与非因果方法,比较确定性、无向和因果三种条件化策略,发现因果方法能提升数据集增强的下游性能与公平性,为数据生成协议设计提供指导。

AI 中文摘要

医学图像分析常常受到有偏数据集的阻碍,这可能导致模型产生偏差并限制其临床适用性。缓解此类偏差的一种有前景的策略是用合成图像扩充训练数据。反事实(CF)生成就是这样一种策略,尽管该术语有两种不同的含义:在一些工作中,反事实是通过基于结构因果模型的因果干预产生的,而在另一些工作中,反事实是通过非因果的图像编辑或常规条件生成模型产生的,例如改变解剖结构或添加病理。在本研究中,我们探讨了这一区别,并评估了其对医学图像增强的实际影响。我们比较了三种条件化策略:\textit{确定性}(Deterministic),即在保持其余变量固定的同时改变所选变量;\textit{无向}(Undirected),即根据学习到的统计关联更新变量而不指定因果方向;以及\textit{因果}(Causal),即沿有向因果图传播干预。我们分析了这些选择如何影响生成的图像,并探讨了基于因果的方法何时能改善数据集增强或带来有限的益处。特别是,我们评估了下游性能和公平性,其中公平性指在敏感子群体中对数据集偏差的敏感性降低。我们的实验表明,使用因果方法生成合成训练数据可以带来切实的益处,这些见解为机器学习从业者有效设计数据生成协议提供了宝贵的指导。

英文摘要

Medical image analysis is often hindered by biased datasets, which can lead to biased models and limited clinical applicability. A promising strategy for mitigating such biases is to augment training data with synthetic images. Counterfactual (CF) generation is one such strategy, though the term is used in two different senses: in some works, CFs are produced through causality-based interventions derived from structural causal models, whereas in others, they are produced by non-causal image edits or conventional conditional generative models, such as altering anatomy or adding pathologies. In this work, we study this distinction and evaluate its practical consequences for medical image augmentation. We compare three conditioning strategies: $\textit{Deterministic}$, which changes selected variables while holding the remaining variables fixed; $\textit{Undirected}$, which updates variables according to learned statistical associations without assigning causal directions; and $\textit{Causal}$, which propagates interventions along a directed causal graph. We analyse how these choices affect the resulting images, and explore when causally grounded methods improve dataset augmentation or bring limited benefit. In particular, we assess downstream performance and fairness, where fairness refers to reduced sensitivity to dataset biases across sensitive subgroups. Our experiments demonstrate that using a causal approach to synthetic training data generation can lead to tangible benefits, with these insights offering valuable guidance to machine learning practitioners for the effective design of data generation protocols.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑