基准测试GPT-5在生物医学自然语言处理中的应用
Benchmarking GPT-5 for biomedical natural language processing
- University of Minnesota(明尼苏达大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过统一基准评估GPT-5在生物医学NLP任务中的性能,发现其优于GPT-4o,尤其在推理密集型问答上,且成本更低,并提出了分层提示策略以平衡精度与效率。
AI中文摘要:
生物医学文献和临床叙述对自然语言理解提出了多方面的挑战,从精确的实体提取和文档综合到多步骤的诊断推理。本研究扩展了一个统一基准,在零样本、单样本和五样本提示下评估GPT-5和GPT-4o在五个核心生物医学NLP任务上的表现:命名实体识别、关系提取、多标签文档分类、摘要和简化,以及九个扩展的生物医学问答数据集,涵盖事实知识、临床推理和多模态视觉理解。使用标准化提示、固定解码参数和一致的推理流程,我们根据官方定价评估了模型性能、延迟和令牌归一化成本。GPT-5持续优于GPT-4o,在推理密集型数据集(如MedXpertQA和DiagnosisArena)上提升最大,在多模态问答中也有稳定改进。在核心任务中,GPT-5在化学命名实体识别和ChemProt得分上表现更好,但在疾病命名实体识别和摘要方面仍低于领域调优的基线。尽管GPT-5生成了更长的输出,但其延迟相当,每次正确预测的有效成本降低了30%至50%。细粒度分析显示在诊断、治疗和推理子类型上有所改进,而边界敏感的提取和证据密集的摘要仍然具有挑战性。总体而言,GPT-5接近部署就绪的生物医学问答性能,同时在准确性、可解释性和经济效率之间提供了有利的平衡。结果支持分层提示策略:对于大规模或成本敏感的应用采用直接提示,对于分析复杂或高风险场景采用思维链支架,突显了在精度和事实保真度至关重要的场景中持续需要混合解决方案。
英文摘要:
Biomedical literature and clinical narratives pose multifaceted challenges for natural language understanding, from precise entity extraction and document synthesis to multi-step diagnostic reasoning. This study extends a unified benchmark to evaluate GPT-5 and GPT-4o under zero-, one-, and five-shot prompting across five core biomedical NLP tasks: named entity recognition, relation extraction, multi-label document classification, summarization, and simplification, and nine expanded biomedical QA datasets covering factual knowledge, clinical reasoning, and multimodal visual understanding. Using standardized prompts, fixed decoding parameters, and consistent inference pipelines, we assessed model performance, latency, and token-normalized cost under official pricing. GPT-5 consistently outperformed GPT-4o, with the largest gains on reasoning-intensive datasets such as MedXpertQA and DiagnosisArena and stable improvements in multimodal QA. In core tasks, GPT-5 achieved better chemical NER and ChemProt scores but remained below domain-tuned baselines for disease NER and summarization. Despite producing longer outputs, GPT-5 showed comparable latency and 30 to 50 percent lower effective cost per correct prediction. Fine-grained analyses revealed improvements in diagnosis, treatment, and reasoning subtypes, whereas boundary-sensitive extraction and evidence-dense summarization remain challenging. Overall, GPT-5 approaches deployment-ready performance for biomedical QA while offering a favorable balance of accuracy, interpretability, and economic efficiency. The results support a tiered prompting strategy: direct prompting for large-scale or cost-sensitive applications, and chain-of-thought scaffolds for analytically complex or high-stakes scenarios, highlighting the continued need for hybrid solutions where precision and factual fidelity are critical.