AI 中文总结
JustLLMGRPO仅对LLM提示词策略应用GRPO,在冻结Sana生成器的前提下,提升了胸部X射线生成的质量,同时维持了提示词对齐度,实现了最优性能。
AI 中文摘要
文本条件下的胸部X射线(CXR)生成旨在合成能忠实呈现指定病变的真实放射影像。现有工作主要通过更新图像生成器来提升质量,在完成CXR领域适配后默认将提示词视为固定不变。本文表明,这种以生成器为中心的视角忽略了一个巨大的优化维度。在冻结经CXR适配的Sana生成器的前提下,使用未修改的大语言模型(LLM)进行一次提示词重新表述,可将RadDINO-FID从54.225降至27.572。提示词分析显示,该LLM会抑制时间对比、不确定性及其他不可渲染的报告内容,同时强调可见的放射学病变。然而,无约束的重新表述会使BioViL-T与源提示词的对齐度从0.695降至0.609。因此,本文提出JustLLMGRPO,仅对LLM提示词策略应用标准的组相对策略优化(GRPO),同时保持Sana冻结。组相对的放射学感知图像反馈在保留视觉焦点的同时维持了与源提示词的对齐。在CheXGenBench上,JustLLMGRPO将RadDINO-FID降至26.780,较直接提示提升了50.6%,且保持了对齐度(0.696 vs 0.695),还实现了最先进的分布覆盖和下游分类效用。这些结果表明,向适配后的生成器输入放射学信息的表达方式中仍存在巨大的性能潜力,代码已公开。
英文摘要
Text-conditioned chest X-ray generation aims to synthesize realistic radiographs that faithfully depict specified findings. Existing work has primarily improved quality by updating image generators, implicitly treating prompts as fixed after CXR-domain adaptation. We show that this generator-centric view leaves a substantial optimization dimension underexplored. With a CXR-adapted Sana generator frozen, one-pass reformulation by an unmodified LLM reduces RadDINO-FID from 54.225 to 27.572. Prompt analysis shows that the LLM suppresses temporal comparisons, uncertainty, and other non-renderable report content while emphasizing visible radiographic findings. However, unconstrained reformulation reduces BioViL-T alignment with source prompts from 0.695 to 0.609. We therefore introduce JustLLMGRPO, which applies standard Group Relative Policy Optimization (GRPO) only to the LLM prompt policy while keeping Sana frozen. Group-relative radiology-aware image feedback retains visual focus while preserving source-prompt alignment. On CheXGenBench, JustLLMGRPO reduces RadDINO-FID to 26.780, a 50.6% improvement over direct prompting, while maintaining alignment (0.696 versus 0.695). It also achieves state-of-the-art distribution coverage and downstream classification utility. These results show that substantial performance can remain latent in how radiographic information is expressed to an adapted generator. Code is publicly available at https://github.com/pxcai/JustLLMGRPO.