arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05579cs.CVcs.AIeess.IV

ViT3Flow:用于脊柱侧凸术后X光片合成的测试时训练Transformer MeanFlow

ViT3Flow: A Test-Time Training Transformer MeanFlow for Postoperative Radiograph Synthesis in Scoliosis

Rui Tang, Sicheng Yang, Moxin Zhao, Hongqiu Wang, Guankun Wang, Lei Zhu, Hongliang Ren, Menglin Cong, Nan Meng

首次发表
浏览论文内容

中文总结 AI 辅助

针对脊柱侧凸术后X光片合成,提出ViT3Flow单步MeanFlow框架,结合测试时训练token混合器与DRICA注意力,在ScoliSurg数据集上实现最优图像质量与解剖保真度。

中文摘要 AI 辅助

从术前X光片预测术后脊柱形态可为脊柱侧凸手术规划提供宝贵支持,但由于手术矫正会引起较大的空间变化,同时必须忠实保留解剖结构,这一任务仍具挑战性。我们将该问题定义为术后脊柱侧凸X光片合成,并构建了ScoliSurg,这是该任务的首个配对数据集,包含632对术前-术后全脊柱X光片及结构化形态信息。我们进一步提出ViT$^{3}$Flow,一种单NFE条件MeanFlow框架,用于高效的术后X光片合成。ViT$^{3}$Flow将手术矫正建模为有限区间生成传输,并用测试时训练token混合器替代传统自注意力,该混合器针对每个病例的解剖结构和畸形模式执行样本特定的内部适应。此外,脊柱形态提取代理从术前X光片中提取主弯区域和方向的分布。这些分布引导诊断路由区间交叉注意力(DRICA),该机制从独立的术前token流中执行区间依赖的垂直、水平、联合和全局检索。这一设计使得不断演化的术后表征能够在整个传输过程中整合空间对应的解剖证据。在ScoliSurg上的大量实验表明,ViT$^{3}$Flow在感知图像质量、解剖保真度和临床相关几何精度方面均达到对比方法中的最佳性能,且仅需一次网络评估。这些结果凸显了ViT$^{3}$Flow在脊柱侧凸手术规划中实现高效且解剖保真的术后X光片合成的潜力。

英文摘要

Predicting postoperative spinal morphology from preoperative radiographs could provide valuable support for scoliosis surgical planning, but remains challenging because surgical correction induces large spatial changes while anatomical structures must be faithfully retained. We formulate this problem as postoperative scoliosis radiograph synthesis and construct ScoliSurg, the first paired dataset for this task, comprising 632 preoperative--postoperative whole-spine radiograph pairs with structured morphology information. We further propose ViT$^{3}$Flow, a single-NFE conditional MeanFlow framework for efficient postoperative radiograph synthesis. ViT$^{3}$Flow models surgical correction as finite-interval generative transport and replaces conventional self-attention with test-time-training token mixers that perform sample-specific inner adaptation to the anatomy and deformity pattern of each case. In addition, a Spinal Morphology Extraction Agent extracts distributions of dominant-curve region and direction from the preoperative radiograph. These distributions guide Diagnosis-Routed Interval Cross-Attention (DRICA), which performs interval-dependent vertical, horizontal, joint, and global retrieval from a separate preoperative token stream. This design enables the evolving postoperative representation to incorporate spatially corresponding anatomical evidence throughout the transport process. Extensive experiments on ScoliSurg demonstrate that ViT$^{3}$Flow achieves the best performance among the compared methods in perceptual image quality, anatomical fidelity, and clinically relevant geometric accuracy, while requiring only a single network evaluation. These results highlight the potential of ViT$^{3}$Flow for efficient and anatomically faithful postoperative radiograph synthesis in scoliosis surgical planning.

发表机构

  • The University of Hong Kong(香港大学)
  • The University of Hong Kong-Shenzhen Hospital(香港大学深圳医院)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • The Hong Kong University of Science and Technology(香港科技大学)
  • The Chinese University of Hong Kong(香港中文大学)
  • Qilu Hospital of Shandong University(山东大学齐鲁医院)

机构由 AI 辅助整理,请以论文原文为准。

↑