形式几何中采样LLM推理的诊断:覆盖度、实现度与有效性证据
Diagnosing Sampled LLM Reasoning in Formal Geometry: Coverage, Realization, and Validity Evidence
AI总结:
针对形式几何中采样LLM推理,提出CRV评估协议,区分覆盖度、实现度与有效性证据,实验显示三者独立且应分别报告。
AI中文摘要:
重复采样可以揭示正确的数值答案,但既不产生可靠的系统输出,也不产生有支撑的推导。我们提出了覆盖度、实现度与有效性证据(CRV),一种针对形式几何状态上采样大语言模型(LLM)推理的评估协议。覆盖度是答案可用性,实现度是在冻结候选池上的读出准确性,有效性证据是标签盲审批评者对推导支撑的判断,而非证明证书。CRV在比较读出之前冻结每个候选池,并通过正确答案多重性和问题内区分度分析覆盖的失败案例。在HardShift441上,一个441道问题的集合,参考求解器留下406道未解问题,LoRA适配的Qwen2.5-7B生成器获得24.2%的平均单样本准确率和68.9%的pass@16,而验证器加权自洽性(WSC)达到38.0%。当正确答案在池中仅出现一次或两次时,读出准确性尤其低。在另一次对195个覆盖问题的构建审计中,批评者将12个正确答案代表标记为有支撑,181个为反驳,2个为不确定。这些结果表明,来自批评者的覆盖度、实现度和有效性证据是不同量,应分别报告。
英文摘要:
Repeated sampling can reveal a correct numerical answer without yielding either a reliable system output or a supported derivation. We present Coverage, Realization, and Validity Evidence (CRV), an evaluation protocol for sampled large language model (LLM) reasoning over formal geometry states. Coverage is answer availability, realization is readout accuracy on the frozen candidate pool, and validity evidence is a label-blinded critic judgment of derivational support rather than a proof certificate. CRV freezes each candidate pool before comparing readouts and analyzes covered failures by correct-answer multiplicity and within-problem discrimination. On HardShift441, a 441-problem set for which a reference solver leaves 406 problems unsolved, a LoRA-adapted Qwen2.5-7B generator obtains 24.2% average single-sample accuracy and 68.9% pass@16, whereas verifier-weighted self-consistency (WSC) reaches 38.0%. Readout accuracy is particularly low when the correct answer occurs only once or twice in the pool. In a separate constructed audit of 195 covered problems, the critic labels 12 correct-answer representatives as supported, 181 as refuted, and two as uncertain. These results show that coverage, realization, and validity evidence from the critic are distinct quantities and should be reported separately.