arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估面向纵向临床笔记的基于LLM问答的生物医学重排序

Evaluating Biomedical Reranking for LLM-Based Question Answering over Longitudinal Clinical Notes

Maryam Shahbaz Ali, Laura B. Strachan, Caitlin Sherman, Mark Kovler, Eleanor Mackey, Syed Muhammad Anwar

arXiv 2610.01324首次发表:更新:

发表机构

Children’s National Hospital; University of Florida; George Washington University(国家儿童医院; 佛罗里达大学; 乔治华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估了生物医学重排序在纵向临床笔记的LLM问答中提升证据检索和答案正确性的效果,发现检索增益未完全转化为答案质量提升。

AI 中文摘要

针对特定患者的临床问答需要在冗长且异质的纵向临床记录中定位正确的证据,这些记录中相关事实可能分散在不同就诊中、在复制转抄的笔记中重复出现,或使用不同的临床术语表达。我们评估了在本地部署的用于纵向临床笔记的检索增强生成流程中,生物医学重排序是否能改善证据选择及下游答案质量。该流程结合了PubMedBERT稠密检索、BM25词法检索、加权倒数排名融合以及MedCPT交叉编码器重排序。在来自200名减肥手术患者队列的1000个开放式和封闭式问答对中,重排序将前10项中的精确源块检索率(Hit@10)从46.6%提高到60.6%,平均倒数排名从0.2371提高到0.3252。使用Qwen3-8B生成时,本地评判者评估的答案正确率从44.8%提高到48.6%。这些结果表明,生物医学重排序可以改善有限上下文窗口内相关临床证据的放置,尽管检索增益并未按比例转化为答案正确性的增益。

英文摘要

Patient-specific clinical question answering requires locating the right evidence within long, heterogeneous longitudinal clinical records in which relevant facts may be scattered across encounters, repeated in copied-forward notes, or expressed using different clinical terminology. We evaluated whether biomedical reranking can improve evidence selection and downstream answer quality in a locally deployed retrieval-augmented generation pipeline for longitudinal clinical notes. The pipeline combines PubMedBERT dense retrieval, BM25 lexical retrieval, weighted reciprocal-rank fusion, and MedCPT cross-encoder reranking. Across 1,000 open- and closed-ended question-answer pairs from a cohort of 200 bariatric surgery patients, reranking increased exact source-chunk retrieval within the top 10 items, Hit@10 from 46.6% to 60.6% and mean reciprocal rank from 0.2371 to 0.3252. With Qwen3-8B generation, local judge-assessed answer correctness increased from 44.8% to 48.6%. These results show that biomedical reranking can improve the placement of relevant clinical evidence within a limited context window, although gains in retrieval do not translate proportionally into gains in answer correctness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑