通过轨迹积分反馈调节解剖感知奖励用于体积计算断层扫描分析
Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis
- Zhejiang University(浙江大学)
- DAMO Academy, Alibaba Group(阿里集团达摩院)
- Hupan Lab(虎扑实验室)
- University of Electronic Science and Technology of China(电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出了一种新的框架,通过轨迹积分反馈GRPO(TIF-GRPO)来改进医疗视觉语言模型在三维CT分析中的性能,通过引入临床异常基准评估子系统(CABS)来解决优化目标与临床严谨性之间的不匹配问题,提升异常检测和临床准确性。
AI中文摘要:
医学视觉-语言模型(VLMs)已迅速发展为通用多模态助手,但其在三维计算机断层扫描(CT)分析中的应用仍受到优化目标与临床严谨性之间持续不匹配的限制。当前的强化学习(RL)范式仍然依赖于词汇代理信号,导致``评估幻觉'',即模型优化语言流畅性而非事实性临床正确性,从而导致诊断性关键错误。为弥合这一差距,我们引入了临床异常基准评估子系统(CABS),一个将放射学报告分解为可验证的临床语义单元的结构化系统。利用CABS,我们识别出标准RL中的``机理分歧'',即表面相似性奖励驱动策略梯度绕过医学事实。因此,我们提出了轨迹积分反馈GRPO(TIF-GRPO),一种将控制理论原理整合到策略优化中的新框架。通过将临床推理建模为伪时间轨迹以发现异常,TIF-GRPO通过积分反馈回路调节解剖感知奖励,该回路将持续遗漏视为累积状态误差,并将幻觉视为过度的控制努力。在3D CT基准测试中,我们的方法显著提高了异常检测和临床忠实度,建立了医疗VLMs中细粒度调节的新范式。我们的项目可在GitHub上获取。
英文摘要:
Medical vision-language models (VLMs) have rapidly advanced as general-purpose multimodal assistants, yet their deployment in 3D Computed Tomography (CT) analysis remains constrained by a persistent mismatch between optimization objectives and clinical rigor. Current Reinforcement Learning (RL) paradigms still rely on lexical proxy signals that induce ``\textit{Evaluation Hallucinations}'', where models optimize linguistic fluency rather than factual clinical correctness, leading to diagnostically critical errors. To bridge this gap, we introduce the \textbf{Clinical Abnormality Benchmarking Substrate (CABS)}, a structured system that decomposes radiology reports into verifiable clinical semantic units. Using CABS, we identify a ``\textit{Mechanistic Divergence}'' in standard RL, where surface-similarity rewards drive policy gradients to bypass medical facts. We therefore propose \textbf{Trajectory-Integral Feedback GRPO (TIF-GRPO)}, a novel framework integrating control-theoretic principles into policy optimization. By formulating clinical reasoning as a pseudo-temporal trajectory for anomaly discovery, TIF-GRPO regulates anatomy-aware rewards via an integral feedback loop that penalizes persistent omissions as cumulative state errors and suppresses hallucinations as excessive control effort. Experiments on 3D CT benchmarks demonstrate that our approach significantly enhances abnormality detection and clinical faithfulness, establishing a new paradigm for fine-grained regulation in medical VLMs. Our project is available at \href{https://github.com/ZJU4HealthCare/TIF-GRPO}{GitHub}.