arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越目标分数:测量基于扩散的医学图像编辑中的偏离目标漂移

Beyond Target Scores: Measuring Off-Target Drift in Diffusion-Based Medical Image Editing

Todd Zhou

arXiv 2607.16291首次发表:更新:

AI 中文总结

研究基于扩散的医学图像编辑中偏离目标漂移问题,引入CIB-Med-1基准及受约束扩散引导基线,通过多指标评估编辑效果,实验显示能保留目标进展并减少偏离目标漂移,强调应从轨迹级语义控制角度评估医学图像编辑。

AI 中文摘要

扩散模型现在可以以视觉上合理的方式编辑医学图像,但标准评估问题过于狭窄:目标分数是否增加?在临床成像中,目标发现与合并症、采集效应和选择偏差纠缠在一起,因此模型可能通过改变相关的非目标发现而不是分离预期的病理来显得成功。我们引入了CIB-Med-1,这是一个用于胸部X光片中受控生物标志物编辑的轨迹级基准。CIB-Med-1通过校准的目标进展、反转率和在14个临床相关的干扰轴上的偏离目标语义漂移来评估定向胸腔积液编辑。该基准揭示了一种奖励作弊失败模式,即扩散编辑器在增加积液分数的同时,会改变实质、心肺纵隔、胸膜、慢性或伪影相关的发现。我们还提出了一个受约束的扩散引导基线,该基线在有界的偏离目标变化的情况下优化目标进展。在保留的X光片中,受约束的编辑器保留了目标进展(无约束引导时的趋势ρ=0.88,而无约束引导时为0.90),同时将偏离目标漂移的中位数从0.46降低到0.20,第90百分位数漂移从0.98降低到0.33。漂移幅度跟踪经验性目标与偏离目标的关联,支持语义不稳定性是结构化而非偶然的观点。一项针对放射科实习生的盲法人体验证探针进一步表明,与预期进展顺序的一致性更强(Pix2Pix的τ=0.29, 而此方法的τ=0.61)。这些结果表明,医学图像编辑应作为轨迹级语义控制进行评估,而不是作为终点分数最大化。

英文摘要

Diffusion models can now edit medical images in visually plausible ways, but the standard evaluation question is too narrow: did the target score increase? In clinical imaging, target findings are entangled with co-morbidities, acquisition effects, and selection bias, so a model can appear successful by changing correlated non-target findings rather than isolating the intended pathology. We introduce CIB-Med-1, a trajectory-level benchmark for controlled biomarker editing in chest radiography. CIB-Med-1 evaluates directional pleural effusion editing through calibrated target progression, inversion rate, and off-target semantic drift over 14 clinically motivated nuisance axes. The benchmark exposes a reward-hacking failure mode in which diffusion editors increase effusion scores while simultaneously altering parenchymal, cardiomediastinal, pleural, chronic, or artifact-related findings. We further present a constrained diffusion guidance baseline that optimizes target progression subject to bounded off-target change. Across held-out radiographs, the constrained editor preserves target progression ($ρ_{\mathrm{trend}}=0.88$ vs. $0.90$ for unconstrained guidance) while reducing median off-target drift from $0.46$ to $0.20$ and 90th-percentile drift from $0.98$ to $0.33$. Drift magnitude tracks empirical target--off-target association, supporting the view that semantic instability is structured rather than incidental. A blinded human validation probe with radiology trainees further shows stronger agreement with intended progression orderings ($τ=0.61$ vs.\ $0.29$ for Pix2Pix). These results argue that medical image editing should be evaluated as trajectory-level semantic control, not as endpoint score maximization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑