用于条件医学图像生成的组合奖励模型
Compositional Reward Models for Conditional Medical Image Generation
浏览论文内容
中文总结 AI 辅助
本研究提出PRISM组合奖励模型框架,通过分层约束传播机制优化条件医学图像生成,在三个医学成像数据集上训练下游模型时,多项指标较基准方法均有显著提升。
中文摘要 AI 辅助
获取高质量的带标注医学图像数据对训练深度学习模型至关重要;然而,标注过程成本高昂、耗时且需要领域专业知识。ControlNet等条件扩散模型提供了一种替代方案,可基于语义掩码和文本生成图像。但现有方法无法捕捉领域专家期望的精细属性(如强度和纹理)以及语义一致性,限制了其在下游任务中的有效性。近期尝试通过强化学习微调解决这些问题,但由于依赖单一标量奖励,该奖励将不同类型的失败模式混为一谈,仅提供微弱的校正信号,因此效果有限。我们提出PRISM,一种用于条件医学图像生成的组合奖励模型(Compositional Reward Model, CRM)框架。我们不分配单一奖励,而是将图像质量分解为基于验证器的阶段,每个阶段从精细到粗糙的属性评估不同方面的正确性,包括低级属性(强度和纹理)、与条件输入的结构对齐以及高级语义保真度。这些阶段性奖励通过分层约束传播(Hierarchical Constrained Propagation, HCP)机制组合,该机制强制执行从精细到粗糙的正确性概念,确保在获得高级别奖励前解决低级缺陷,防止较易目标掩盖关键失败。我们在涵盖不同医学成像任务的三个数据集上评估PRISM:PanNuke(多类细胞分割)、CeDeM(绒毛/隐窝检测与测量)和ISIC(皮肤病变分类)。使用PRISM生成的数据训练下游模型,相比最接近的基准方法取得了改进,包括PanNuke上mDice提升2.3%、CeDeM上平均相对误差(Mean Relative Error, MRE)降低8.5%,以及ISIC的F1提升5.9%。
英文摘要
Acquiring high quality annotated medical image data is critical for training deep learning models; however, annotation is expensive, time consuming, and requires domain expertise. Conditional diffusion models, such as ControlNet, offer an alternative by generating images conditioned on semantic masks and text. However, existing approaches fail to capture fine grained properties (e.g., intensity and texture), as well as semantic consistency expected by domain experts, limiting their effectiveness for downstream tasks. Recent attempts to address these issues using reinforcement learning fine-tuning remain limited due to the reliance on a single scalar reward, which conflates diverse failure modes and provides weak corrective signals. We propose PRISM, a Compositional Reward Model (CRM) framework for conditional medical image generation. Instead of assigning a single reward, we decompose image quality into verifier grounded stages, each evaluating a distinct aspect of correctness from fine to coarse properties, including low level attributes (intensity and texture), structural alignment with conditioning inputs, and high level semantic fidelity. These stage wise rewards are composed through a Hierarchical Constrained Propagation (HCP) mechanism that enforces a fine to coarse notion of correctness, ensuring that lower level deficiencies are resolved before higher level rewards are accrued, preventing easier objectives from masking critical failures. We evaluate PRISM across three datasets spanning diverse medical imaging tasks: PanNuke (multi-class cell segmentation), CeDeM (villi/crypt detection and measurement), and ISIC (skin lesion classification). Training downstream models with data generated by PRISM yields improvements over closest baselines, including a 2.3% increase in mDice on PanNuke, a 8.5% reduction in Mean Relative Error (MRE) on CeDeM, and increases ISIC F1 by 5.9%.
发表机构
- Yardi School of Artificial Intelligence(亚尔迪人工智能学院)
- Indian Institute Of Technology, Delhi(德里印度理工学院)
- Indian Institute of Science(印度科学学院)
- Department of Computer Science and Engineering(计算机科学与工程系)
机构由 AI 辅助整理,请以论文原文为准。