发表机构
Nanyang Technological University; Wuhan University(南洋理工大学; 武汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
JUDGESTEALER是首个跨评估协议提取LLM评判能力的高效查询框架,通过利用跨协议一致性等技术提升提取效果,在三类评估任务上优于基线且具鲁棒性。
AI 中文摘要
大语言模型(LLM)评判器正越来越多地应用于各类评估场景,其评判能力成为极具价值的知识产权。然而,黑盒访问特性使得这些能力易受模型提取攻击。现有提取方法未专门针对LLM评判器,且在受限查询预算下对多种评估协议的支持有限。本研究提出JUDGESTEALER,首个查询高效的模型提取框架,用于复现点评分、成对比较、列表排序协议下的评判能力。JUDGESTEALER利用跨协议的强一致性获取点评分,并将其转化为成对和列表监督信号,无需额外查询目标模型;它还基于语义多样性、预测不确定性和潜在评判偏差动态选择点输入,以捕捉有价值的评判模式并提升查询效率;此外,应用评分平滑和多协议校验,保留评分的序数结构,缓解代理模型适配时的灾难性遗忘。对最先进的LLM评判器和奖励模型的大量实验表明,JUDGESTEALER始终优于现有提取基线,在点、成对、列表评估上分别达到最高73.3%、87.0%、71.6%的准确率;该框架在不同代理模型规模、适配策略和推理设置下均保持有效,且对代表性提取防御措施具有鲁棒性。
英文摘要
Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to model extraction attacks. Existing extraction methods do not specifically target LLM judges and provide limited support for multiple evaluation protocols under restricted query budgets. In this study, we propose JUDGESTEALER, the first query-efficient model extraction framework for replicating judging capabilities across pointwise scoring, pairwise comparison, and listwise ranking protocols. JUDGESTEALER exploits the strong cross-protocol agreement to acquire pointwise scores and transform them into pairwise and listwise supervisions without additional victim queries. To capture informative judge patterns and improve query efficiency, JUDGESTEALER dynamically selects pointwise inputs based on semantic diversity, predictive uncertainty, and potential judge biases. It further applies score smoothing and multi-protocol review to preserve the ordinal structure of scores and mitigate catastrophic forgetting during surrogate adaptation. Extensive experiments on state-of-the-art LLM-as-a-judge and reward models show that JUDGESTEALER consistently outperforms existing extraction baselines, achieving up to 73.3%, 87.0%, and 71.6% accuracy for pointwise, pairwise, and listwise evaluation, respectively. JUDGESTEALER also remains effective across different sur- rogate model scales, adaptation strategies, and reasoning settings. Moreover, JUDGESTEALER demonstrates robustness against representative extraction defenses.
Comments20 pages, 8 figures