arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14226cs.CL

语料库特征刻画与逆宪法微调用于风格感知的放射学报告

Corpus Characterization and Inverse Constitutional Fine-Tuning for Style-Aware Radiology Reports

Sarah Y. Li, Elijah Renner, Rayan Ansari, Alaa Youssef

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过语料库特征刻画和逆宪法微调,使放射学报告生成模型在结构和词汇上对齐真实医生风格,显著提升文本匹配分数。

中文摘要 AI 辅助

自动化放射学报告生成在诊断准确性方面取得了快速发展,然而生成的报告在结构、措辞和不确定性语言方面常常偏离真实放射科医生写作的风格惯例,这一差距对临床医生的信任和用户体验有直接影响。为解决这一问题,我们使用Bio-ClinicalBERT嵌入、UMAP降维和HDBSCAN聚类对CheXpert Plus数据集中的2,000份报告进行了风格变异特征刻画,识别出五种不同的报告模式,这些模式在病理关注点、叙事结构和词汇偏好上有所不同。基于这些发现,我们调整了逆宪法AI框架,从放射科医生撰写的报告对中推导出以风格为中心的宪法,无需正式的偏好数据集。该宪法编码了语气、措辞、不确定性校准和报告结构的惯例,并纳入对MedGemma-4B基础模型在25,245个CheXpert Plus训练对上的监督微调中。与未调优的基线相比,宪法微调在文本对齐方面产生了显著提升(BLEU-4:从0.006到0.308;ROUGE-L:从0.171到0.484)。这些提升显示出结构和词汇对齐的质性转变,而非边际改进,因为基线模型由于格式不匹配而产生接近零的分数。总体而言,我们确立了语料库级别的风格特征刻画和宪法建模作为一种有效且数据高效的策略,用于生成符合真实放射科医生写作惯例的放射学报告。

英文摘要

Automated radiology report generation has advanced rapidly in diagnostic accuracy, yet generated reports frequently diverge from the stylistic conventions of authentic radiologist writing in structure, diction, and uncertainty language, a gap which has direct implications for clinician trust and user experience. To address this, we characterize stylistic variation across 2,000 reports from the CheXpert Plus dataset using Bio-ClinicalBERT embeddings, UMAP dimensionality reduction, and HDBSCAN clustering, identifying five distinct reporting patterns differing in pathology focus, narrative structure, and lexical preference. Drawing on these findings, we adapt the inverse constitutional AI framework to derive a style-focused constitution from radiologist-written report pairs without requiring a formal preference dataset. This constitution, encoding conventions of tone, diction, uncertainty calibration, and report structure, is incorporated into the supervised fine-tuning of a MedGemma-4B base model on 25,245 CheXpert Plus training pairs. Constitutional fine-tuning produces a substantial increases in text alignment (BLEU-4: 0.006 to 0.308; ROUGE-L: 0.171 to 0.484) relative to the untuned baseline. These gains show a qualitative shift in structural and lexical alignment rather than marginal improvement, as the baseline model produces near-zero scores due to format mismatch. Overall, we establish corpus-level style characterization and constitutional modeling as an effective and data-efficient strategy for producing radiology reports that conform to authentic radiologist writing conventions.

发表机构

  • McLean High School(麦克莱恩高中)
  • Stanford University(斯坦福大学)
  • Stanford Cardiovascular Institute(斯坦福心血管研究所)
  • Stanford School of Medicine(斯坦福医学院)

机构由 AI 辅助整理,请以论文原文为准。

↑