arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05125cs.LG

基于Logit Bias的单查询黑盒校准审计

Single-Query Black-Box Calibration Auditing via Logit Bias

Roman Plaud, Antoine Saillenfest, Matthieu Labeau, Thomas Bonald, Willem Waegeman

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM API隐藏输出概率的问题,利用logit_bias参数提出单查询审计方法,构建可证一致的二元任务真实校准误差估计量,实现高效黑盒基础模型校准审计。

中文摘要 AI 辅助

评估大语言模型(LLMs)的校准度对其作为零样本分类器的安全部署至关重要。但商业API提供商日益隐藏标准校准指标所需的连续输出概率。为绕过这种不透明性,我们证明任何暴露logit_bias参数的LLM API都可通过数学操作,对每个样本仅用一次查询就能评估精确概率阈值。利用该机制,我们提出了一种新颖且可证一致的二元任务真实校准误差估计量,为黑盒基础模型的审计提供了高效框架。

英文摘要

Evaluating the calibration of Large Language Models (LLMs) is critical for their safe deployment as zero-shot classifiers. Yet, commercial API providers increasingly hide the continuous output probabilities required by standard calibration metrics. To bypass this opacity, we demonstrate that any LLM API exposing a logit\_bias parameter can be mathematically manipulated to evaluate exact probability thresholds using strictly one query per sample. Leveraging this mechanism, we introduce a novel and provably consistent estimator of the True Calibration Error for binary tasks. Our approach therefore provides an efficient framework for auditing black-box foundation models.

发表机构

  • Institut Polytechnique de Paris(巴黎理工学院)
  • Onepoint(万珀因特公司)
  • Ghent University(根特大学)

机构由 AI 辅助整理,请以论文原文为准。

↑