arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于廉价辅助信号的语言模型条件评估

Conditional Evaluation of Language Models with Cheap Auxiliary Signals

Zhi Zhang, Lingfeng Lyu, Yue Kang, Doudou Zhou

arXiv 2608.16210首次发表:更新:

发表机构

University of California, Los Angeles; University of Science and Technology of China; Microsoft; National University of Singapore(加利福尼亚大学洛杉矶分校; 中国科学技术大学; 微软公司; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有语言模型条件评估中廉价辅助信号存在偏差的问题,提出半监督估计器LACE,结合局部中心化与岭控制变量,在多个基准数据集上完成实证评估。

AI 中文摘要

聚合准确率无法体现模型的成功与失败之处。仅通过黄金标签估计条件性能剖面成本高昂,而诸如LLM评判分数、成对比较、置信度分数及评判分歧特征等廉价辅助信号可针对每个基准项收集,但往往存在偏差或校准不当。我们提出LACE(Local Augmented Control-Variate Evaluation,局部增强控制变量评估),这是一种用于条件大语言模型(LLM)评估的半监督估计器。关键步骤为局部中心化:在目标剖面区域内减去廉价信号的条件均值后,任何线性增强项的条件均值均为零,因此无法改变待估计量。增强系数仅用于提升效率,局部岭控制变量结合了标注子集的黄金标签残差均值与全项池的廉价信号均值。我们证明了该校准自由识别、分组剖面的无偏性、中心化线性增强内的局部神谕最优性,以及对估计系数的一阶适应性。所得增益公式由总体局部R²支配,其刻画了从廉价信号可获得的效率随剖面值的变化情况。我们还推导了直接成对模型差距及部署加权分数的对应估计器。我们在MATH-500、ScienceQA、MMLU、WinoGrande、HellaSwag、TruthfulQA、GSM8K和ARC上对主要性能剖面估计器进行了实证评估。

英文摘要

Aggregate accuracy hides where models succeed and fail. Estimating conditional performance profiles from gold labels alone is expensive, while cheap auxiliary signals such as LLM-judge scores, pairwise comparisons, confidence scores, and judge-disagreement features can be collected for every benchmark item but are often biased or miscalibrated. We propose LACE (Local Augmented Control-Variate Evaluation), a semi-supervised estimator for conditional LLM evaluation. The key step is local centering: after subtracting the conditional mean of a cheap signal within the target profile region, any linear augmentation has zero conditional mean and therefore cannot change the estimand. The augmentation coefficient is used only for efficiency, and a local ridge control variate combines a gold-label residual mean from the labeled subset with a cheap-signal mean from the full item pool. We prove calibration-free identification, unbiasedness for grouped profiles, local oracle optimality within centered linear augmentations, and first-order adaptivity to the estimated coefficient. The resulting gain formula is governed by a population local $R^2$, which characterizes how the efficiency attainable from the cheap signals varies across profile values. We also derive corresponding estimators for direct paired model gaps and deployment-weighted scores. We empirically evaluate the primary performance-profile estimator on MATH-500, ScienceQA, MMLU, WinoGrande, HellaSwag, TruthfulQA, GSM8K, and ARC.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑