arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Claude Fable 5在生物医学挑战问题上的能力

Capabilities of Claude Fable 5 on Biomedical Challenge Problems

Dominic Okonkwo, Magnus Hodgson, Temitope I. David, Susan Adanna Ihejirika

arXiv 2607.10849首次发表:更新:

发表机构

School of Computing, University of Georgia; Department of Chemistry, University of Illinois; Institute of Bioinformatics, University of Georgia(佐治亚大学计算学院; 伊利诺伊大学化学系; 佐治亚大学生物信息学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究评估Claude Fable 5在八个生物医学基准上的能力,以确定性评分针对固定答案键评估,含两个Claude前身和GPT-5作基线。发现Fable 5拒绝率因基准而异,排除拒绝项后准确率超其他模型,还识别出两种拒绝模式,其生物医学应用主要受限在参与意愿。

AI 中文摘要

前沿语言模型越来越多地在生物医学基准上进行评估,但大多数已发表的评估存在两个问题:传统基准接近饱和,开放式回答由其他语言模型评分。我们使用确定性评分,针对固定答案键,在八个生物医学基准(四个文本和四个多模态)上评估了Anthropic最强大的公开可用模型Claude Fable 5,包括两个Claude前身和GPT-5作为基线。在每个结果表中,拒绝被作为一个不同的结果进行跟踪。这一决定产生了论文的核心发现。根据基准不同,Fable 5拒绝8.0%至99.4%的问题,这一模式在其前身和GPT-5中均不存在。排除被拒绝的项目后,Fable 5在本研究的每个基准上的准确率超过或达到其他模型。我们识别出两种可区分的拒绝模式:一种集中在MedQA和MedXpertQA MM的基础科学和机制内容上,另一种在RareBench上是单独的疾病领域模式,其中先天性代谢疾病表现几乎被普遍拒绝,而成人自身免疫性疾病表现则不然。Fable 5在生物医学应用上的主要限制是参与意愿,而非能力。

英文摘要

Frontier language models are increasingly evaluated on biomedical benchmarks, but two problems undermine most published evaluations: legacy benchmarks are near-saturated, and open-ended responses are graded by other language models. We evaluate Claude Fable 5, Anthropic's most capable publicly available model, across eight biomedical benchmarks, four text and four multimodal, using deterministic scoring against fixed answer keys throughout. We include two Claude predecessors and GPT-5 as baselines. Refusal is tracked as a distinct outcome in every result table. That decision produces the paper's central finding. Fable 5 refuses between 8.0% and 99.4% of questions depending on the benchmark, a pattern absent in both predecessors and in GPT-5. Once refused items are excluded from the denominator, Fable 5's accuracy exceeds or meets every other model on every benchmark in this study. We identify two distinguishable refusal patterns: one concentrating in basic-science and mechanism content across MedQA and MedXpertQA MM, confirmed independently on two benchmarks using each benchmark's own category labels; and a separate disease-domain pattern on RareBench, where inborn metabolic disease presentations are refused near-universally while adult-onset autoimmune presentations are not. The primary constraint on Fable 5's biomedical usefulness is willingness to engage, not capability once it does.

Comments15 pages, 6 tables, 4 figures, appendix with qualitative examples

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑