基于Logit Bias的单查询黑盒校准审计
Single-Query Black-Box Calibration Auditing via Logit Bias
浏览论文内容
中文总结 AI 辅助
该研究针对LLM API隐藏输出概率的问题,利用logit_bias参数提出单查询审计方法,构建可证一致的二元任务真实校准误差估计量,实现高效黑盒基础模型校准审计。
中文摘要 AI 辅助
评估大语言模型(LLMs)的校准度对其作为零样本分类器的安全部署至关重要。但商业API提供商日益隐藏标准校准指标所需的连续输出概率。为绕过这种不透明性,我们证明任何暴露logit_bias参数的LLM API都可通过数学操作,对每个样本仅用一次查询就能评估精确概率阈值。利用该机制,我们提出了一种新颖且可证一致的二元任务真实校准误差估计量,为黑盒基础模型的审计提供了高效框架。
英文摘要
Evaluating the calibration of Large Language Models (LLMs) is critical for their safe deployment as zero-shot classifiers. Yet, commercial API providers increasingly hide the continuous output probabilities required by standard calibration metrics. To bypass this opacity, we demonstrate that any LLM API exposing a logit\_bias parameter can be mathematically manipulated to evaluate exact probability thresholds using strictly one query per sample. Leveraging this mechanism, we introduce a novel and provably consistent estimator of the True Calibration Error for binary tasks. Our approach therefore provides an efficient framework for auditing black-box foundation models.
发表机构
- Institut Polytechnique de Paris(巴黎理工学院)
- Onepoint(万珀因特公司)
- Ghent University(根特大学)
机构由 AI 辅助整理,请以论文原文为准。