arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VoxStruct3D:用于体素空间3D MRI合成的结构引导流匹配方法

VoxStruct3D: Structure-Leading Flow Matching for Voxel-Space 3D MRI Synthesis

Fang Li, Yang Gao, Shihao Zou, Weixin Si, Hongyu Wu, Qing Xia, Shuai Li, Aimin Hao

arXiv 2608.04557首次发表:更新:

AI 中文总结

本文提出VoxStruct3D,一种体素空间流匹配框架,通过结构优先策略引导,在T1加权脑MRI合成任务上实现了特征分布对齐、样本多样性等指标的最优性能,生成的体积解剖连贯且视觉真实。

AI 中文摘要

高保真3D MRI合成既需要全局连贯的解剖结构,也需要细粒度的体素级细节。尽管潜在扩散使体积生成变得可行,但其图像自编码器会引入重建瓶颈,可能限制最终体积中可恢复的精细细节。我们提出VoxStruct3D,这是一种体素空间流匹配框架,采用干净数据预测目标直接建模全分辨率MRI体积。其体素体生成器(Volumetric Voxel Generator, VVG)结合了因式分解的3D patch嵌入、重叠上采样、时间调制残差细化和跳跃融合,使相邻token能联合重建共享体素区域并抑制patch边界伪影。为通过显式解剖先验补充直接体素空间建模,我们进一步提出结构优先、图像跟随(Structure-First, Image-Follows, SFIF)策略:冻结的预训练3D医学编码器和StructVAE提取保留主要解剖结构的紧凑结构token,结构引导调度使这些token的轨迹超前于图像轨迹;Patch-Aligned RoPE实现不等token网格的空间对齐,非对称注意力强制执行从结构到图像的单向引导。在病理和健康T1加权脑MRI数据集上的实验表明,VoxStruct3D在特征分布对齐、样本多样性和感知质量方面实现了最强的整体性能,生成解剖连贯且视觉真实的体积。

英文摘要

High-fidelity 3D MRI synthesis requires both globally coherent anatomy and fine-grained voxel-level detail. Although latent diffusion makes volumetric generation tractable, its image autoencoder introduces a reconstruction bottleneck that can limit the fine detail recoverable in the final volume. We present VoxStruct3D, a voxel-space flow-matching framework that directly models full-resolution MRI volumes using a clean-data prediction objective. Its Volumetric Voxel Generator (VVG) combines factorized 3D patch embedding with overlapping upsampling, time-modulated residual refinement, and skip fusion, enabling neighboring tokens to jointly reconstruct shared voxel regions and suppress patch-boundary artifacts. To complement direct voxel-space modeling with an explicit anatomical prior, we further introduce a Structure-First, Image-Follows (SFIF) strategy. A frozen pretrained 3D medical encoder and a StructVAE extract compact structure tokens that preserve dominant anatomy, while a structure-leading schedule keeps their trajectory ahead of the image trajectory. Patch-Aligned RoPE spatially aligns the unequal token grids, and asymmetric attention enforces one-way guidance from structure to image. Experiments on pathological and healthy T1-weighted brain MRI datasets show that VoxStruct3D achieves the strongest overall performance across feature-distribution alignment, sample diversity, and perceptual quality, producing anatomically coherent and visually realistic volumes.

CommentsProject page: https://neesky.github.io/VoxStruct3D/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑