校准小型语言模型以进行声明的核查价值检测
Calibrating Small Language Models for Claim Check-Worthiness Detection
浏览论文内容
中文总结 AI 辅助
针对小型语言模型(SLMs)核查声明价值时的准确性与成本问题,提出轻量级校准方法NN-PPI,可使SLMs达到大型语言模型的准确性,大幅降低大规模核查的成本。
中文摘要 AI 辅助
评估声明的核查价值是自动化事实核查流程中的关键第一步。本研究的动机源于一家初创公司早期部署时面临的实际挑战:对每一条传入声明都运行大型语言模型(LLM)会产生过高的成本和延迟,而较小的模型则会牺牲准确性。我们提出了NN-PPI,这是Prediction-Powered Inference(PPI)的逐点扩展方法,作为一个轻量级的事后层,在推理时校准模型预测,无需重新训练底层模型。NN-PPI根据基线模型的规模和性能,实现了12%至33.80%的加权F1值提升,使小型语言模型(SLMs)的表现与更大的LLM相当。除了少样本SLMs外,NN-PPI还进一步改进了一个已投入生产部署的微调模型,证明了残差校准与监督微调具有互补性。通过从成本低一个数量级的模型中恢复LLM级别的准确性,它大幅降低了大规模运行准确的核查价值检测的成本。我们的代码和数据可在此处的URL获取。
英文摘要
Assessing claim check-worthiness is an essential first step in automated fact-checking pipelines. This work is motivated by a real deployment challenge at an early-stage startup: running large language models (LLMs) over every incoming claim is cost- and latency-prohibitive, yet smaller models sacrifice accuracy. We propose NN-PPI, a pointwise extension of Prediction-Powered Inference (PPI) that calibrates model predictions at inference time as a lightweight post-hoc layer, without re-training the underlying model. NN-PPI achieves weighted F1 gains ranging from 12% to 33.80% depending on the size and performance of the baseline model, bringing SLMs on par with larger LLMs. Beyond few-shot SLMs, NN-PPI further improves a production-deployed fine-tuned model, demonstrating that residual calibration is complementary to supervised fine-tuning. By recovering LLM-level accuracy from models that are an order of magnitude cheaper to serve, it makes accurate check-worthiness detection substantially cheaper to operate at scale. Our code and data can be found at https://anonymous.4open.science/r/arr-claim-worthiness-F237.
发表机构
- Factiverse AS(法克蒂弗斯公司)
- University of Stavanger(斯塔万格大学)
- Stockholm University(斯德哥尔摩大学)
机构由 AI 辅助整理,请以论文原文为准。