大型语言模型(LLMs)能否判断优质假说?基于logit的能量评分在科学假说排序上优于提示式LLM评判器
Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking
浏览论文内容
中文总结 AI 辅助
本研究针对科学假说评估难题,提出基于logit的能量评分方法,在12学科1323篇论文的基准测试中,该方法的Hit@1显著优于提示式LLM评判器,展现出模型内在置信度的应用潜力。
中文摘要 AI 辅助
大型语言模型(LLMs)正越来越多地被用于科学假说生成,但对生成的假说进行评估仍是可信赖的AI驱动科学工作流面临的挑战。现有方法常将LLMs用作评判器或依赖语义相似度,这类方法可能更青睐熟悉的想法而非新颖的想法。我们提出一种基于logit的能量评分方法,利用语言模型的内在置信度而非比较判断来评估假说。我们在12个学科的1323篇论文上对7种语言模型进行了基准测试,每篇论文搭配其假说和15个不正确的替代选项。内在评分在两个评分器上汇总的Hit@1达到33.0%,而提示式列表排序的Hit@1为16.6%。最强的配置是采用基于logit的能量评分的10亿参数模型,其Hit@1达到53.1%,不过这是事后选择的14种模型-评分器组合中的最大值。总体而言,模型内在置信度在科学假说评估方面展现出潜力,本研究也为可信赖的AI驱动科学发现的基于置信度方法的未来研究提供了动力。
英文摘要
Large language models (LLMs) are increasingly used for scientific hypothesis generation. However, evaluating generated hypotheses remains a challenge for trustworthy AI-enabled scientific workflows. Existing approaches often use LLMs as judges or rely on semantic similarity, which can favor familiar ideas over novel ones. We propose a logit-based energy scoring method that evaluates hypotheses using a language model's intrinsic confidence rather than comparative judgment. We benchmarked seven language models on 1,323 papers across 12 disciplines. Each paper was paired with its hypothesis and fifteen incorrect alternatives. Intrinsic scoring reached 33.0% Hit@1 pooled across both scorers, compared with 16.6% for prompted listwise ranking. The strongest configuration, a 1-billion-parameter model using logit-based energy scoring, reached 53.1%, though this was the maximum across 14 model-by-scorer combinations selected post hoc. Overall, intrinsic model confidence shows potential for scientific hypothesis evaluation. This study also motivates future research on confidence-based methods for trustworthy AI-enabled scientific discovery.
发表机构
- Oak Ridge National Laboratory(橡树岭国家实验室)
机构由 AI 辅助整理,请以论文原文为准。