CORDIAL:从少量标签校准有序大语言模型输出
CORDIAL: Calibrating Ordinal LLM Outputs from Few Labels
浏览论文内容
中文总结 AI 辅助
提出CORDIAL,通过五参数噪声通道校正LLM有序输出,少量标签即可实现高精度校准,并在多种LLM和数据集上优于现有方法。
中文摘要 AI 辅助
大语言模型(LLM)可以将文本转化为在有序标度上的分布,但该分布是一种有噪声的测量:饱和、压缩或夸大,并沿一致方向产生偏差。我们提出CORDIAL,该方法将模型输出视为真实标签的有噪声读数,并通过包含五个可解释参数的通道对其进行校正。该通道足够小,其后验可从少量标签中平均得到,我们证明所得校准保持一阶随机优势。在Amazon评论和CMU-MOSEI转录文本上,使用四种LLM,CORDIAL在80个设置中的76个(标签数量为5至100)中,在九个校准器中取得最低对数损失;使用20个标签和主要的7B读取器时,其性能与使用28-54个标签的最强基线相当。相同的后验使我们能从其他任务学习先验并融合多个LLM。无限制校准器(如Dirichlet校准)仅在校准集增长至数百或数千时才能超越它。
英文摘要
A large language model (LLM) can turn a text into a distribution over an ordered scale, but that distribution is a noisy measurement: saturated, compressed or exaggerated, and biased in a consistent direction. We propose CORDIAL, which treats the model's output as a noisy reading of the true label and corrects it with a channel of five interpretable parameters. The channel is small enough for its posterior to be averaged from a handful of labels, and we prove that the resulting calibration preserves first-order stochastic order. On Amazon reviews and CMU-MOSEI transcripts with four LLMs, CORDIAL has the lowest log loss among nine calibrators in 76 of 80 settings with 5 to 100 labels; with 20 labels and the main 7B reader, it matches the strongest baseline using 28-54 labels. The same posterior lets us learn priors from other tasks and fuse several LLMs. Unrestricted calibrators such as Dirichlet calibration overtake it only as the calibration set grows into the hundreds or thousands.
发表机构
- The University of Melbourne(墨尔本大学)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。