arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2512.12868cs.CLcs.AIcs.LG

计数线索:一个轻量级概率基线可以匹配LLM

Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM

  • Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

Furong Jia, Yuan Pu, Finn Guo, Monica Agrawal

更新

AI总结:

本文提出了一种轻量级概率排名器FBPR,通过共现统计信息实现与LLM相当的性能,展示了概率基线在临床诊断中的互补优势。

AI中文摘要:

大型语言模型(LLM)在多项选择的临床诊断基准测试中表现出色,但不清楚这种表现有多大程度上反映了底层的概率推理。我们通过MedQA中的问题进行研究,该任务是选择最可能的诊断。我们引入了基于频率的概率排名器(FBPR),这是一种轻量级方法,通过在大规模语料库中对概念-诊断共现统计信息进行平滑的朴素贝叶斯方法来对选项进行评分。当共现统计信息来自OLMo和Llama的预训练语料库时,FBPR的性能与相应预训练在相同语料库上的LLM相当。直接LLM推理和FBPR在大部分问题上得到正确答案,正确答案的重叠部分仅略高于随机猜测,表明每种方法都有互补的优势。这些发现突显了显式概率基线的持续价值:它们提供了一个有意义的性能参考点,并为潜在的混合系统提供了补充信号。虽然LLM的性能似乎由一种不同于简单频率聚合的机制驱动,但我们的研究表明,一种类似于历史上扎根、低复杂度专家系统的方法仍然在基准性能中占有一席之地。

英文摘要:

Large language models (LLMs) excel on multiple-choice clinical diagnosis benchmarks, yet it is unclear how much of this performance reflects underlying probabilistic reasoning. We study this through questions from MedQA, where the task is to select the most likely diagnosis. We introduce the Frequency-Based Probabilistic Ranker (FBPR), a lightweight method that scores options with a smoothed Naive Bayes over concept-diagnosis co-occurrence statistics from a large corpus. When co-occurrence statistics were sourced from the pretraining corpora for OLMo and Llama, FBPR achieves comparable performance to the corresponding LLMs pretrained on that same corpus. Direct LLM inference and FBPR largely get different questions correct, with an overlap only slightly above random chance, indicating complementary strengths of each method. These findings highlight the continued value of explicit probabilistic baselines: they provide a meaningful performance reference point and a complementary signal for potential hybridization. While the performance of LLMs seems to be driven by a mechanism other than simple frequency aggregation, we show that an approach similar to the historically grounded, low-complexity expert systems still accounts for a substantial portion of benchmark performance.

↑