arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向多模态临床诊断的排序感知提示优化

Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis

Tian Xia, Minghao Liu, Yiqing Liang, Laixi Shi, Jiayun Wang

arXiv 2609.40361首次发表:更新:

发表机构

Harvard University; University of California, Santa Cruz; Brown University; Johns Hopkins University; Georgia Institute of Technology(哈佛大学; 加州大学圣克鲁兹分校; 布朗大学; 约翰斯·霍普金斯大学; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态临床诊断中类别不平衡问题,提出成对排序帕累托提示演化(Ranking-PE),用AUROC替代准确率优化提示,在MIMIC上显著提升排序性能。

AI 中文摘要

多模态大语言模型(MLLMs)正在快速推进临床诊断,但其适配流程仍锚定于基于准确率的目标。临床数据高度类别不平衡:一个常数多数类预测器可达到90%以上的准确率,却在临床上毫无用处。因此,我们评估并优化AUROC,这是一种无阈值的评分,将正类排在负类之上,且对类别平衡不变。我们聚焦于MLLMs中的提示优化。诸如GEPA之类的反思性方法使用二元评分矩阵,每行对应一个评估实例,每列对应一个候选提示;单元格记录每个实例的正确性,因此列平均值即为准确率,并驱动候选选择。我们引入成对级别的帕累托提示演化(Ranking-PE),将每个正确性行替换为在(正类,负类)实例对上的成对排序行:若候选对正类的评分高于配对的负类,则单元格为1。列平均值随后等于经验AUROC(依据Wilcoxon-Mann-Whitney恒等式)。我们在提示演化搜索所读取的所有三个层面——决定帕累托支配的评分矩阵、对反思LM的逐示例反馈以及最终候选选择——应用此替换,且无需额外模型调用,也无需替代损失。在MIMIC上针对三种疾病,基于准确率的提示演化可能降低排序性能;Ranking-PE逆转了这一点,在微调的Qwen3-VL-8B上比基于准确率的方法高出+5.8 AUROC百分点,在MedGemma-4B上高出+16.2个百分点。消融实验检查了每个设计组件,并表明医学级视觉骨干——通过视觉编码器微调的SFT或医学预训练——是提示搜索无法替代的先决条件;我们的方法将反思性提示演化从纯文本数据扩展到多模态临床决策。

英文摘要

Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based objectives. Clinical data are heavily class-imbalanced: a constant-majority predictor can score above 90% accuracy while being clinically useless. We therefore evaluate and optimize for AUROC, a threshold-free score that ranks positives above negatives and is invariant to class balance. We focus on prompt optimization in MLLMs. Reflective methods such as GEPA use a binary scores matrix with one row per evaluation instance and one column per candidate prompt; cells record per-instance correctness, so the column average is accuracy and drives candidate selection. We introduce pair-level Pareto prompt evolution (Ranking-PE), which replaces each correctness row with a pairwise-ordering row over (positive, negative) instance pairs: the cell is 1 if the candidate scores the positive higher than the paired negative. The column average then equals empirical AUROC (by the Wilcoxon-Mann-Whitney identity). We apply this swap at all three layers the prompt evolution search reads from - the scores matrix that decides Pareto dominance, the per-example feedback to the reflection LM, and final candidate selection - at no extra model calls and with no surrogate loss. Across three diseases on MIMIC, accuracy-based prompt evolution can degrade ranking; Ranking-PE reverses this, beating the accuracy-based recipe by +5.8 AUROC pp on fine-tuned Qwen3-VL-8B and +16.2 pp on MedGemma-4B. Ablations examine each design component and show that a medical-grade visual backbone - via vision-encoder-tuned SFT or medical pretraining - is a prerequisite that prompt search cannot replace - our recipe extends reflective prompt evolution from text-only data to multimodal clinical decision-making.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑