arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Jev在医学中的应用:基准评估。初步结果

Jev in Medicine: A Benchmark Evaluation

Alfredo Madrid-García, Beatriz Merino-Barbancho

arXiv 2609.34024首次发表:更新:

AI 中文总结

本研究评估了非生成式模型Jev在四个医学基准上的表现,发现其准确率在研究摘要任务上与前沿LLM相当,但在复杂诊断病例上显著较低,且速度快、成本极低,需任务特定验证后才能用于临床。

AI 中文摘要

Jev是一种非生成式“系统一”模型,它为预定义的答案选项分配概率,并且无法回答这些选项之外的问题。其在医学问答和基于病例的诊断推理任务上的准确性和校准情况尚不清楚。我们在四个医学基准上评估了Jev 1.13:MetaMedQA、PubMedQA、DiagnosisArena-MCQ和NEJM病例挑战。GPT-6 Sol(使用中等推理和不使用推理)作为参考。主要结果是top-1准确率;关键的次要结果是校准、选择性预测和对不可回答问题的识别。所有8,469个请求均返回了有效答案。在PubMedQA上,Jev的准确率与使用中等推理的GPT-6 Sol相当(78.4%对78.2%),在MetaMedQA上较低(74.8%对82.7%),在DiagnosisArena-MCQ(59.8%对82.4%)和NEJM病例(61.8%对82.4%)上则低得多。在MetaMedQA上,Jev的概率校准最佳(预期校准误差0.063对0.146),其概率至少为0.9的答案(占问题的52.9%)准确率为93.4%,但GPT-6 Sol在接受类似比例问题时的准确率也相同。在DiagnosisArena-MCQ上,Jev的概率区分能力较差(AUROC 0.645对0.768)。在162个正确答案为“我不知道或无法回答”的问题中,Jev选择了该选项的占10.5%(GPT-6 Sol为8.6%)。中位延迟为0.27-0.31秒;全部2,823个项目花费0.08美元。Jev快速且廉价,其准确率在研究摘要上接近前沿LLM,但在考试题上较低,在复杂诊断病例上则低得多。在临床使用前需要进行特定任务的验证。

英文摘要

Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. All 8,469 requests returned a valid answer. Jev's accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%;), lower on MetaMedQA (74.8% vs 82.7%) and much lower on DiagnosisArena-MCQ (59.8% vs 82.4%;) and the NEJM cases (61.8% vs 82.4%). On MetaMedQA, Jev's probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability of at least 0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev's probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was "I don't know or cannot answer", Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27-0.31 s; all 2,823 items cost USD 0.08. Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Task-specific validation is required before clinical use.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑