From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
从分数到步骤:诊断和改进证据医学计算中LLM的性能
机构 * Department of Computer Science, Yale University, CT, USA(耶鲁大学计算机科学系) ; Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA贝福德医疗中心健康组织与实施研究中心) ; Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(UMass洛厄尔矿尔计算机与信息科学学院) ; Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(UMass阿默斯特马宁信息与计算机科学学院)
专题命中 RAG评测 :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI
AI总结 本文提出MedRaC框架,通过分步评估和代码执行提升LLM在证据医学计算中的准确性,揭示现有评估方法的不足,并推动临床可信度的提升。
Comments Equal contribution for the first two authors. To appear as an Oral presentation in the proceedings of the Main Conference on Empirical Methods in Natural Language Processing (EMNLP) 2025