MedForge-RSI:通过递归自我改进实现医学深度伪造检测
MedForge-RSI: Medical Deepfake Detection via Recursive Self-Improvement
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对医学深度伪造检测在部署偏移下性能下降的问题,提出MedForge-RSI递归自我改进框架,在冻结模型权重下通过20轮自主分析错误和开发工具,将平均准确率从75.0%提升至84.4%。
AI中文摘要:
文本引导的图像编辑器可以生成高保真度的医学深度伪造图像,对临床影像的可靠性构成挑战。尽管基于推理的检测器在分布内表现强劲,但在部署偏移下性能会大幅下降。MedForge-Reasoner是一个通过监督微调和强化学习训练的80亿参数视觉语言模型,在其目标分布上达到了99.2%的准确率,却将40%的真实扫描误分类,在未见过的生成器上仅达到77%的准确率,并在传输失真下降至59%。通过传统重训练来适配此类模型成本高昂,需要大规模监督和专家设计的准则。我们提出了MedForge-RSI,一种递归自我改进框架,使已部署的检测器能够在保持模型权重冻结的情况下进行适配。在20轮迭代中,检测器分析已验证的错误,积累可复用的经验,并自主开发图像分析工具,同时独立的验收测试仅保留经过验证的改进。在49个注册配置中,MedForge-RSI将四个测试集上的平均准确率从75.0%提升至84.4%。在保留的4000张图像评估中,干净图像的准确率从76.5%提升至87.9%,传输失真图像的准确率从59.0%提升至70.5%,其中真实图像召回率的提升最为显著。对所有49条轨迹的对照分析识别出哪些自我改进机制在不同随机种子间可复现,表明验收测试是最大的独立贡献者,并揭示了失败适配的分类。我们发布了完整的轨迹,包括所有被拒绝和回滚的更改。
英文摘要:
Text-guided image editors can generate high-fidelity medical deepfakes, challenging the reliability of clinical imagery. Although reasoning-based detectors perform strongly in distribution, they degrade substantially under deployment shift. MedForge-Reasoner, an 8B vision-language model trained with supervised fine-tuning and reinforcement learning, achieves 99.2% accuracy on its target distribution, yet misclassifies 40% of authentic scans, reaches only 77% accuracy on unseen generators, and falls to 59% under transmission distortion. Adapting such models through conventional retraining is costly, requiring large-scale supervision and expert-designed guidelines. We introduce MedForge-RSI, a recursive self-improvement framework that enables a deployed detector to adapt while keeping its model weights frozen. Over 20 rounds, the detector analyzes verified errors, accumulates reusable experience, and autonomously develops image-analysis tools, while an independent acceptance test retains only validated improvements. Across 49 registered configurations, MedForge-RSI increases average accuracy over four test sets from 75.0% to 84.4%. On a held-out 4,000-image evaluation, it improves clean accuracy from 76.5% to 87.9% and transmission-distorted accuracy from 59.0% to 70.5%, with the largest gains in authentic-image recall. Controlled analysis across all 49 trajectories identifies which self-improvement mechanisms replicate across seeds, shows acceptance testing to be the largest individual contributor, and reveals a taxonomy of failed adaptations. We release the complete trajectories, including all rejected and rolled-back changes.