arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MS-Exam-Gen:用于评估LLM在多发性硬化MRI文本知识上的源接地基准构建

MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge

Abdul Basit, Muhammad Abdullah Hanif, Muhammad Shafique

arXiv 2610.06170首次发表:更新:

发表机构

New York University Abu Dhabi (NYUAD)(纽约大学阿布扎比分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MS-Exam-Gen提出一个可复现框架,从66个来源构建含3,058个多选题的MS-MRI知识基准,通过自动审计评估12个LLM,准确率范围46.9%-89.7%,支持项目级和主题级评估。

AI 中文摘要

生物医学大语言模型(LLM)评估需要对狭窄、不断演变且基于来源的子专业知识进行可审计的评估。多发性硬化MRI(MS-MRI)提供了一个高风险文本知识测试案例,因为正确的推理需要最新的诊断标准、标准化采集与报告知识、纵向监测概念、病灶形态学以及识别困难的模拟病变。我们提出了MS-Exam-Gen,一个用于构建和审计基于文本的多选题(MCQ)基准的可复现框架,用于MS-MRI知识;它不评估直接的MRI图像解读。MS-Exam-Gen针对源接地标准、协议、报告和鉴别诊断。该框架结合了专家来源索引、面向考试的主题归纳、基于证据的MCQ生成、自动化质量审计、同族一致性筛查和经验校准。从包含66个来源的语料库索引为4,289个检索块,该流程生成了一个包含3,058个项目的锁定候选基准,涵盖16个主题和53个子主题。在12个主要LLM端点上的评估产生了36,696个项目级预测,并在42.8个百分点的准确率范围(89.7%至46.9%)内区分了性能。在这些端点中,25.5%的项目被至少四个端点遗漏。生成后审计显示,刷新构建减少了可测量的答案线索,而选项顺序测试表明绝对MCQ分数仍对位置敏感。生成的构建标签仍是元数据,而非经过验证的心理测量类别。由于专家裁决和完全选项顺序平衡仍是未来工作,MS-Exam-Gen不是临床认证考试。它应被解释为自动过滤的、基于来源的候选基准和可复现的审计工作流,用于项目级和主题特定的LLM评估。

英文摘要

Biomedical large language model (LLM) evaluation requires auditable assessment of narrow, evolving, source-grounded subspecialty knowledge. Multiple sclerosis MRI (MS-MRI) provides a high-stakes textual-knowledge test case because correct reasoning requires current diagnostic criteria, standardized acquisition and reporting knowledge, longitudinal monitoring concepts, lesion morphology, and recognition of difficult mimics. We present MS-Exam-Gen, a reproducible framework for constructing and auditing a text-based multiple-choice question (MCQ) benchmark for MS-MRI knowledge; it does not evaluate direct MRI image interpretation. MS-Exam-Gen targets source-grounded criteria, protocols, reporting, and differential diagnosis. The framework combines expert-source indexing, exam-oriented topic induction, evidence-grounded MCQ generation, automated quality audits, a same-family consistency screen, and empirical calibration. From a 66-source corpus indexed into 4,289 retrieval chunks, the pipeline produced a locked 3,058-item candidate benchmark spanning 16 topics and 53 subtopics. Evaluation across 12 primary LLM endpoints yielded 36,696 item-level predictions and separated performance over a 42.8-percentage-point accuracy range (89.7% to 46.9%). Across these endpoints, 25.5% of items were missed by at least four. Post-generation audits showed that refreshed construction reduced measurable answer cues, while option-order testing showed that absolute MCQ scores remain position-sensitive. Generated construction labels remain metadata rather than validated psychometric categories. Because expert adjudication and full option-order counterbalancing remain future work, MS-Exam-Gen is not a clinically certified examination. It should be interpreted as an automatically filtered, source-grounded candidate benchmark and reproducible audit workflow for item-level and topic-specific LLM evaluation.

Comments7 pages, 3 figures. Accepted for publication to BHI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑