arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22824physics.med-phcs.AIcs.CV

用于CT重建的智能自动研究

Agentic Autoresearch for CT Reconstruction

Andreas Maier, Lucas Kachelriess, Siming Bayer, Yixing Huang, Yan Xia, Amber Simpson, Moritz Zaiss

首次发表
浏览论文内容

中文总结 AI 辅助

研究CT重建方法比较难题,利用大语言模型智能体构建智能循环,独立实现、调整并基准测试26种方法,重组为紧凑求解器,发现理想数据排名无法预测现实噪声下表现,噪声会颠倒排名,重新训练可恢复部分排名,强调基准测试需考虑多种现实因素。

中文摘要 AI 辅助

公平比较CT重建方法既费力又大多依赖人工,且许多基准测试使用理想化数据。我们探究大语言模型(LLM)智能体能否自行开展重建研究工作,以及基于理想数据的排名能否预测现实噪声下的表现。我们构建了一个智能循环:智能体编辑求解器、运行短集群作业、读取一个固定指标并进行修正。该指标是视野内相对于FBP基线的校准余量分数,所有方法共享相同的可微扇束投影仪。我们在梅奥低剂量CT(噪声受限)和无噪声DL - 稀疏视图挑战中的128视图稀疏视图乳腺任务上对26种方法进行基准测试,在验证选定的迭代次数下在留出的测试集上评分。然后在无重新训练的情况下对每个训练好的乳腺模型在有噪声输入(I_0 = 10^5光子)上重新评分,并在匹配噪声下单独重新训练。智能体独立实现、调整并对所有26种方法进行基准测试,并将它们重组为一个969参数的紧凑求解器,该求解器以冠军参数的0.4%在1%水平上与梅奥顶级方法相当。基准测试给出了一组统计上无法区分的顶级方法,而非单一获胜者。轻微的输入噪声几乎使乳腺排名颠倒:无噪声冠军(一种监督图像去噪器,hr 0.89)降至0.00,而一种学习的原始对偶方法升至冠军(从0.72升至0.93)。因此,理想数据排行榜无法预测鲁棒性。这种颠倒属于转移效应,而非永久性缺陷:在匹配噪声下重新训练可恢复大部分干净排名(斯皮尔曼相关系数从0.04升至0.61)。噪声只是众多开放式混杂因素(束硬化、散射、解剖结构、疾病)中最容易的一个,所以没有单一因素挑战能证明通用性。基准测试应同时对广泛的现实因素进行建模。

英文摘要

Comparing CT reconstruction methods fairly is labor-intensive and largely manual, and many benchmarks use idealized data. We ask whether a large language model (LLM) agent can do the labor of reconstruction research on its own, and whether a ranking measured on ideal data predicts behavior under realistic noise. We built an agentic loop: the agent edits a solver, runs a short cluster job, reads one frozen metric, and revises. The metric is a calibrated headroom score against the FBP baseline, inside the field of view; every method shares the same differentiable fan-beam projector. We benchmarked 26 methods on Mayo low-dose CT (noise-limited) and a 128-view sparse-view breast task from the noiseless DL-Sparse-View Challenge, with validation-selected iterations scored on a held-out test set. Every trained breast model was then re-scored on noisy inputs (I_0 = 10^5 photons) without retraining, and separately retrained on matched noise. The agent independently implemented, tuned, and benchmarked all 26 methods, and recombined them into a compact solver of 969 parameters that ties the top Mayo tier at the 1% level using 0.4% of the champion's parameters. Benchmarking gives a tier of statistically indistinguishable top methods, not one winner. Mild input noise nearly inverts the breast ranking: the noiseless champion (a supervised image denoiser, hr 0.89) collapses to 0.00, while a learned primal-dual method rises to champion (0.72 to 0.93). An ideal-data leaderboard therefore does not predict robustness. The inversion is a transfer effect, not a permanent deficit: retraining on matched noise restores much of the clean ranking (Spearman rho 0.04 to 0.61). Noise is only the easiest confounder in an open-ended set (beam hardening, scatter, anatomy, disease), so no single-factor challenge certifies generality. Benchmarks should model a broad spectrum of realistic factors at once.

发表机构

  • Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU)(弗里德里希-亚历山大-埃尔兰根-纽伦堡大学(FAU))
  • Universitätsklinikum Erlangen (UKER)(埃尔兰根大学医院(UKER))
  • Peking University(北京大学)
  • University of Alberta(阿尔伯塔大学)
  • Department Artificial Intelligence in Biomedical Engineering (AIBE), Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU)(弗里德里希-亚历山大-埃尔兰根-纽伦堡大学生物医学工程人工智能系(AIBE))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑