AI 中文总结
研究针对医学图像超分辨率难题,提出MedDiT4SR三流自适应框架,集成多种表示于扩散Transformer块,引入SR Adapter和SA Refiner,实验验证该框架能有效让预训练DiT模型适应不同医学成像域的超分辨率任务。
AI 中文摘要
医学图像超分辨率(MedSR)需要从退化观测中恢复精细解剖结构,同时避免生成先验引入的无根据细节。大规模预训练多模态扩散Transformer提供了强大的视觉先验,但使其适应MedSR并非易事。在传统ControlNet风格的自适应中,低分辨率(LR)图像作为外部条件处理并通过单向连接注入去噪流,导致LR解剖证据无法与不断演变的去噪和语义表示联合更新。我们提出了MedDiT4SR,一个三流自适应框架,将LR、噪声潜在和文本表示集成到相同的多模态扩散Transformer块中。为补充全局令牌交互,引入了超分辨率适配器(SR Adapter)聚合尺度相关的局部令牌并抑制插值引起的冗余。还提出了语义对齐细化器(SA Refiner)使用提示条件语义信息校准局部LR响应。在域内和模态内跨数据集设置下的实验证明了将大规模预训练DiT模型适应不同成像域的医学图像超分辨率的有效性。
英文摘要
Medical image super-resolution (MedSR) requires recovering fine anatomical structures from degraded observations while avoiding unsupported details introduced by generative priors. Large-scale pre-trained multimodal diffusion transformers provide strong visual priors, but their adaptation to MedSR remains non-trivial. In conventional ControlNet-style adaptation, the low-resolution (LR) image is processed as an external condition and injected into the denoising stream through one-way connections. Consequently, LR anatomical evidence cannot be jointly updated with the evolving denoising and semantic representations. We propose MedDiT4SR, a tri-stream adaptation framework that integrates the LR, noisy latent, and text representations into the same multimodal diffusion-transformer blocks. To complement global token interaction, we introduce a Super-Resolution Adapter (SR Adapter) that aggregates scale-dependent local tokens and suppresses interpolation-induced redundancy. We further propose a Semantic Alignment Refiner (SA Refiner) that calibrates local LR responses using prompt-conditioned semantic information. Experiments under both in-domain and within-modality cross-dataset settings demonstrate the effectiveness of adapting large-scale pre-trained DiT models to medical image super-resolution across diverse imaging domains.