arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单公共因子模型下锚定裁判误差相关性的闭式估计器与诊断电池

A Closed-Form Estimator and Diagnostic Battery for Anchor-Judge Error Correlation, Under a Single-Common-Factor Model

Veerendra Kumar Sunkavalli

arXiv 2609.08826首次发表:更新:

AI 中文总结

提出单公共因子模型下锚定裁判误差相关性的闭式估计器,以诊断电池处理不可检验假设,并给出识别条件与验证结果。

AI 中文摘要

当使用外部参考集(锚)将LLM裁判组的误差分解为质量信号和共享共模误差时,标准做法假设锚是无污染的:其误差与裁判的共享误差不相关。我们研究何时可以放弃该假设并用估计值替代。在单公共因子模型下,≥2个裁判和≥2个锚以闭式形式点识别质量方差、共模方差以及每个锚的污染相关性ρ_k,并具有精确的每锚对失败边界;相比之下,指定的清洁锚估计器在其信任的锚本身受污染时,会将受污染的伴随锚报告为完全清洁。由于单公共因子假设本身不可检验,该估计器在交付时带有校准的诊断电池(裁判协方差离散度;过度识别;基于裁判元数据的族块检验,以及去除族级共享残差偏差的族块估计器)、具有实测覆盖率的自助置信区间和弱识别筛查。一个命题映射了哪些违规会使ρ_k产生偏差、偏差方向以及哪些违规可逃避检测。对于有序分数,我们展示了识别层级:当所有变量为有序时,ρ_k在任何锚数量下均不可识别;当裁判为有序且存在≥3个连续锚时则可识别,我们为此情况给出了估计器。在真实数据上,验证是不对称的,我们坦率说明:诊断在拒绝方向上得到验证(我们测试的两个真实面板均被模型充分性预检验正确拒绝),而估计器在模拟中得到验证并在半合成数据上经受压力测试(使用预言校准);尚无真实面板通过预检验,而预检验的存在正是为了说明这一点。所有结果均从随附的带校验和工件离线重放。

英文摘要

When an external reference set (an anchor) is used to decompose an LLM-judge panel's error into a quality signal and a shared common-mode error, standard practice assumes the anchor is uncontaminated: its error uncorrelated with the judges' shared error. We study when that assumption can be dropped and replaced by an estimate. Under a single-common-factor model, >=2 judges and >=2 anchors point-identify the quality variance, the common-mode variance, and each anchor's contamination correlation rho_k in closed form, with an exact per-anchor-pair failure boundary; a designated clean-anchor estimator, by contrast, reports a contaminated companion anchor as fully clean once its trusted anchor is itself contaminated. Because the single-common-factor assumption is itself untestable, the estimator ships gated behind a calibrated diagnostic battery (judge-covariance dispersion; over-identification; a family-block test from judge metadata, with a family-blocked estimator that removes family-level shared-residual bias exactly), bootstrap confidence intervals with measured coverage, and a weak-identification screen. A proposition maps which violations bias rho_k, in which direction, and which evade detection. For ordinal scores we show an identification hierarchy: with all variables ordinal, rho_k is not identified at any number of anchors; with ordinal judges and >=3 continuous anchors it is, and we give an estimator for that case. On real data the validation is asymmetric, and we say so plainly: the diagnostics are validated in the rejecting direction (both real panels we test are correctly rejected by the model-adequacy pre-test), while the estimator is validated in simulation and stress-tested semi-synthetically under oracle calibration; no real panel has yet passed the pre-test, and the pre-test exists precisely to say so. All results replay offline from shipped, checksummed artifacts.

Comments11 pages. The complete reproduction artifact (code, frozen data, one-command replay) accompanies the paper's journal submission as supplementary material, currently under peer review; a public repository link will be added upon publication

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑