arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28538cs.CVcs.LG

ScaFE:基于大语言模型生成的临床特征程序实现数据高效的瘢痕分类

ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs

Ruman Wang, Hangting Ye

AI总结:

ScaFE将LLM临床知识转化为本地可执行的特征程序,结合轻量级随机森林,在瘢痕分类任务中实现数据高效、跨站点的高性能,且决策可审计,优于现有基线模型。

AI中文摘要:

从临床照片中对病理性瘢痕进行分类,需要区分瘢痕疙瘩与增生性瘢痕,却面临专家标注数据有限、不同医院间采集差异大的问题。端到端图像模型仍依赖大量数据,而将照片上传至托管的视觉语言模型(VLM)可能违反本地数据治理要求,且其决策难以复现和审计。本文提出ScaFE(瘢痕特征工程),将大语言模型(LLM)中的临床知识转化为确定性、可执行的特征程序,而非要求模型直接诊断图像。联网的LLM检索临床证据并合成程序,用于测量瘢痕的视觉可评估属性;候选程序在受限本地环境中执行,仅返回聚合验证统计量和特征级SHAP摘要用于迭代修复与优化,原始图像和患者级输出均保留在本地。随后,轻量级随机森林对得到的结构化表示进行分类。在来自三家医院的600张照片上采用留一医院评估,ScaFE实现81.0%的医院宏平均平衡准确率,比最强基线BiomedCLIP高出10.0个百分点;仅用10%的开发数据时,ScaFE仍保持72.0%的平衡准确率,领先幅度达11.8个百分点。迭代优化还将可执行程序比例从66.7%提升至95.0%,最终特征中有91.7%得到验证证据支持。这些结果表明,LLM的知识可通过本地、可审计的特征程序支持数据高效的跨站点医学图像分类,而非依赖VLM的直接决策。

英文摘要:

Classifying pathological scars from clinical photographs requires distinguishing keloids from hypertrophic scars despite limited expert-labeled data and substantial acquisition variation across hospitals. End-to-end image models remain data-dependent, whereas sending photographs to a hosted vision-language model (VLM) may conflict with local data-governance requirements and yields decisions that are difficult to reproduce and audit. We introduce ScaFE (Scar Feature Engineering), which transfers clinical knowledge from a large language model (LLM) into deterministic, executable feature programs instead of asking the model to diagnose images. A web-enabled LLM retrieves clinical evidence and synthesizes programs that measure visually assessable scar attributes. Candidate programs execute in a restricted local environment, and only aggregate validation statistics and feature-level SHAP summaries are returned for iterative repair and refinement; raw images and patient-level outputs remain local. A lightweight Random Forest then operates on the resulting structured representation. On 600 photographs from three hospitals under leave-one-site-out evaluation, ScaFE achieves 81.0% site-macro balanced accuracy, exceeding the strongest baseline, BiomedCLIP, by 10.0 percentage points. With only 10% of the development data, ScaFE retains 72.0% balanced accuracy and an 11.8-point lead. Iterative refinement also raises the executable-program rate from 66.7% to 95.0%, with verified evidence for 91.7% of the final features. These results show that LLM knowledge can support data-efficient, cross-site medical image classification through local and auditable feature programs rather than direct VLM decisions.

↑