发表机构
School of Statistics and Data Science, Southwestern University of Finance and Economics(西南财经大学统计与数据科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出协议级可识别性审计方法,无需调用模型即可验证LLM评估协议的有效性,发现基础准确率与选择性响应保真度存在显著差异,还确定了冻结策略类的最小识别支持。
AI 中文摘要
即使观测协议无法识别其旨在测量的行为属性,大语言模型基准分数仍可能精确。在受控的、基于求解器的设置中,我们针对有限行为策略类形式化了协议级可识别性审计:给定策略H、观测支持O和估计量τ,我们测试O是否能区分每对具有不同τ的策略。该审计无需调用任何模型即可解决我们的诊断案例:仅基础观测将7个冻结确定性策略归为一个等价类;全支持产生7个类且无跨估计量冲突;每个留一法支持都保留了一个构造性冲突见证。实证上,两种约束生成变体的配对有效性均为1.0,但基础准确率与选择性响应保真度存在差异——在6个平衡的 oracle 转移方向上分别为0.620和0.324(聚类自助法95%置信区间[0.600,0.642]与[0.304,0.345]),且在第二个确定性源上再次出现差距(0.646对0.331)。该审计还为冻结策略类综合出最小识别支持O*:2个单元而非完整的36单元张量。此案例表明,评估设计的有效性可在模型推理前通过结构检查,以及为何基础正确性不决定干预响应保真度。
英文摘要
LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. In a controlled, solver-grounded setting, we formalize a protocol-level identifiability audit over a finite behavioral policy class: given policies H, observation support O, and estimand $τ$, we test whether O separates every pair with different $τ$. The audit requires zero model calls and resolves our diagnostic case: base-only observation collapses seven frozen deterministic policies into one equivalence class; full support yields seven classes and no cross-estimand collisions; every leave-one-out support retains a constructive collision witness. Empirically, both constrained-generation variants have pair-validity 1.0, yet base accuracy and selective-response fidelity diverge - 0.620 versus 0.324 across six balanced oracle-transition directions (cluster-bootstrap 95% CI [0.600, 0.642] vs. [0.304, 0.345]) - and the gap recurs on a second deterministic source (0.646 vs. 0.331). The audit also synthesizes a minimum identifying support $O^*$ for the frozen policy class: two cells instead of the full 36-cell tensor. This case shows how evaluation-design validity can be checked structurally before model inference and why base correctness does not determine intervention-response fidelity.
Comments15 pages, 9 figures. Ning Huang, Ziqi Sha, and Wenxuan Tang contributed equally as second authors. Wei Deng is the corresponding author