似然排序在LLM中不像提示那样随规模扩展
Likelihood Ranking doesn't Scale Like Prompting in LLMs
查看机构详情
- University of Pisa(比萨大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究比较了LLM在陈述句似然排序与提示回答两种评估方式下的表现,发现前者随规模增长保持稳定,后者显著提升,表明两者不可互换。
中文摘要 AI 辅助
LLM评估通常通过提示模型生成答案或使用基于似然的指标对候选输出进行评分来进行。然而,在多选题问答中,标准的基于似然的评分仍然以问题和答案集为条件,因此可以利用与提示相同的任务条件答案选择接口。我们研究了一种基于从相同问答对构建的陈述句的似然排序的互补协议。在95个仅解码器模型(参数范围从0.1B到104B)和10个MCQA数据集上,我们发现陈述句似然排序与提示回答之间存在系统性分歧。陈述句似然准确率在不同规模下保持相对稳定,而提示回答则随着规模和指令调优而显著提高。这些结果表明,对受控陈述性备选项的似然偏好与任务条件答案选择探测了模型行为的不同方面,不应被视为可互换的。
英文摘要
LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative statements constructed from the same question--answer pairs. Across 95 decoder-only models, ranging from 0.1B to 104B parameters, and 10 MCQA datasets, we find a systematic divergence between declarative-statement likelihood ranking and prompted answering. Statement-likelihood accuracy remains comparatively stable across scale, whereas prompted answering improves sharply with scale and instruction-tuning. These results suggest that likelihood preferences over controlled declarative alternatives and task-conditioned answer selection probe distinct aspects of model behavior, and should not be treated as interchangeable.