arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31629cs.CLcs.AIcs.LG

ChestPheNoT:从放射学报告中提取可部署、可审计的标签-状态-证据

ChestPheNoT: Deployable, Auditable Label-Status-Evidence Extraction from Radiology Reports

Kai Yu, Chenyu Zhu, Zaifu Zhan, Meijia Song, Min Zeng, Xiaoyi Chen, Mingquan Lin, Rui Zhang

AI总结:

提出CHESTPHENOT,一个0.5-3B的紧凑语言模型,联合提取放射学报告中的发现标签、三类状态和证据片段,通过混合监督和GRPO优化,在跨机构检测中超越CheXbert,实现可审计的本地部署。

AI中文摘要:

从放射学报告中提取结构化表型支持队列构建、质量审计和临床分析,但实际部署需要本地推理和可审计的预测,而专家注释仍然稀缺。传统的标注器提供结构化发现和断言状态,但不提供支持证据,而API托管的大型语言模型在临床文本不能离开机构基础设施时可能不适用。我们提出了CHESTPHENOT,一个紧凑的0.5-3B语言模型,联合提取发现标签、三类状态(存在/不存在/不确定)和逐字支持证据片段。CHESTPHENOT使用混合CheXbert+72B银监督训练,随后进行监督微调和轻量级GRPO优化。在三个涵盖分布内、跨分类和跨机构评估的人工标注金标准集上,3B模型在分布内仍低于其CheXbert银教师,但在分布偏移下具有竞争力,在跨机构检测上显著超过CheXbert(+2.0 F1)。任务特定训练还使3B模型在大多数检测和状态比较中匹配或超过规模大得多的提示模型。对于证据基础提取,超过99%的最终证据片段可在源报告中定位,3B模型达到47.5可审计F1,比Qwen2.5-7B单次提示高出7.6分,接近Qwen2.5-72B。这些结果表明,本地部署模型可以提供竞争性和直接可审计的放射学报告提取,而无需依赖外部推理API。代码和完整的提取/判断提示将在https URL上提供。

英文摘要:

Structured phenotype extraction from radiology reports supports cohort construction, quality auditing, and clinical analytics, but practical deployment requires local inference and auditable predictions, while expert annotations remain scarce. Conventional labelers provide structured findings and assertion states but no supporting evidence, while API-hosted large language models may be unsuitable when clinical text cannot leave institutional infrastructure. We present CHESTPHENOT, a compact 0.5-3B language model that jointly extracts finding labels, three-class status (present/absent/uncertain), and verbatim supporting evidence spans. CHESTPHENOT is trained using hybrid CheXbert+72B silver supervision followed by supervised fine-tuning and lightweight GRPO refinement. Across three human-annotated gold sets spanning in-distribution, cross-taxonomy, and cross-institution evaluation, the 3B model remains below its CheXbert silver teacher in distribution but is competitive under distribution shift, significantly surpassing CheXbert on cross-institution detection (+2.0 F1). Task-specific training also enables the 3B model to match or exceed substantially larger prompted models on most detection and status comparisons. For evidence-grounded extraction, over 99% of final evidence spans are locatable in the source report, and the 3B model achieves 47.5 auditable-F1, outperforming Qwen2.5-7B one-shot prompting by 7.6 points and approaching Qwen2.5-72B. These results demonstrate that locally deployable models can provide competitive and directly auditable radiology-report extraction without relying on external inference APIs. Code and the full extraction/judge prompts will be made available at https://github.com/yukkai/ChestPheNoT.

补充信息

↑