发表机构
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对基于语言模型的评判器成本高、延迟大等问题,提出用程序蒸馏方法,将其决策逻辑提炼成程序委员会。引入PAJAMA系统,实验显示程序评判器性能可匹配大语言模型,还能提高准确率和吞吐量,产生廉价有效奖励信号。
AI 中文摘要
基于语言模型的评判器已成为自动评估的标准,但存在成本高、延迟大及决策不透明等问题,影响其可扩展性和可靠性。本文提出用程序蒸馏解决这些问题,将语言模型的决策逻辑提炼成直接对候选者评分的程序委员会。在此基础上引入PAJAMA系统,合成程序作为评判器,汇总决策形成联合裁决,并纳入回退机制。实验表明程序评判器能匹配13B规模语言模型评判器的性能,PAJAMA提高了准确率和吞吐量,程序评判器还能产生廉价有效的奖励信号。
英文摘要
LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, inference latency, and opaque decisions---limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at evaluation time, we distill its decision logic into a committee of programs that can score candidates directly. These programmatic judges offer transparency, are easily inspected or edited, and eliminate per-sample API costs. Building on this notion, we introduce PAJAMA, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM. Across five datasets and eight model families, we show that programmatic judges match the performance of a 13B-size LLM judge at 47x higher throughput. When using program outputs as routing signals, PAJAMA improves both accuracy and throughput and advances the Pareto frontier. Beyond evaluation, programmatic judges produce cheap and effective reward signals: on RewardBench, a reward model distilled from programs' verdicts outperforms one trained on a proprietary LLM's labels at two orders of magnitude lower API cost.
CommentsProject page and our demonstration can be found in https://sprocketlab.github.io/PAJAMA/