不匹配不代表错误:不完整参考集可反转开放式心智理论追踪中的校准排名
Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking
AI总结:
该研究发现开放式ToM追踪中不完整参考集生成的代理标签可反转模型校准排名,提出TriSource-Restore方法修复置信度并维持排名方向。
AI中文摘要:
开放式心智理论(ToM)追踪器会输出有限参考集中不存在的有效信念。有限参考集加匹配器的流水线会将不匹配的输出标记为错误,从而生成代理标签,这可能反转固定输出下的真分数模型选择。在保持259个信念及配对分数固定的情况下,参考编码将加权流行率从0.783降至0.295,并反转严格真的Brier风险:在参考标签下,冻结的源先验规则使原生置信度领先0.227,而在盲审裁决下则落后0.152,且在全部6种已编写场景中均如此。仅参考的Platt校准器会进一步反转。在发布的含301个问题的NQ-open DPR-BERT流水线中,出现了ICE特定的反转:其平均置信度基线在精确匹配下将实例级校准误差提升0.045,但在人类正确性下将其降低0.074,且两个区间均排除零。在独立编写的OpenToM叙事中,90%-96%经审计的不匹配信念在字面上为真,且配对方向再次反转。精确分解将失真归因于遗漏的真相,且闭式准则可正确分类12个已发布系统的比较结果。冻结审计的回顾性重放显示,50次尝试注释以至少0.996的概率恢复排名方向。TriSource-Restore以概率采样的人类试点为锚,锚定全帧参考标签和冻结自动判断,维持至少名义覆盖率,缩小区间,并针对基率部署门修复置信度。
英文摘要:
Open-ended Theory-of-Mind (ToM) trackers emit valid beliefs absent from finite references. A finite-reference-plus-matcher pipeline marks unmatched outputs false, creating proxy labels that can reverse proper-score model selection on fixed outputs. Holding 259 beliefs and paired scores fixed, reference recoding lowers weighted prevalence from 0.783 to 0.295 and reverses strictly proper Brier risk: a frozen source-prior rule leads native confidence by 0.227 under reference labels and trails by 0.152 under blinded adjudication, in all six authored scenarios. A reference-only Platt recalibrator reverses further. An ICE-specific reversal appears in a released 301-question NQ-open DPR-BERT pipeline: its average-confidence baseline improves instance-level calibration error by 0.045 under exact match but worsens it by 0.074 under human correctness, with both intervals excluding zero. On independently authored OpenToM narratives, 90-96% of audited unmatched beliefs are literally true and the paired direction again reverses. An exact decomposition attributes the distortion to omitted truths, and a closed-form criterion correctly classifies comparisons from twelve released systems. Frozen-audit retrospective replay shows 50 attempted annotations recover ranking direction with probability at least 0.996. TriSource-Restore anchors full-frame reference labels and frozen automatic judgments to a probability-sampled human pilot, maintains at least nominal coverage, narrows intervals, and repairs confidence subject to a base-rate deployment gate.