发表机构
Independent Researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在地面真值延迟等情况下对模型生成战略路线的评估,提出RouteCast机制,通过模型提出候选路线等产生临时预测排名,后续结果评估预测,在回顾性试点中显示出一定区分能力,建立了可行性结果并揭示失败模式。
AI 中文摘要
许多对模型输出的评估要么依赖于在评估时可检查的合同,要么依赖于在操作循环中到达的反馈。我们研究了一种补充设置,其中地面真值被延迟、审查或保密,因此确定性代码无法在评分时检查正确性,而必须发布代码拥有的临时预测。RouteCast为模型生成的类型化战略路线实例化了这种机制:模型提出候选路线和结构化因素;时间点证据、参考类和确定性转换产生临时预测排名;后续结果评估预测。在对21个二元结果案例(6个阳性,15个阴性)的回顾性风险试点中,全数据包RouteCast分数显示出初步的回顾性区分(AUC 0.756,95% CI [0.471,0.980]),而盲目大语言模型判断达到AUC 0.678 [0.419,0.897],身份暴露大语言模型判断达到AUC 0.761 [0.515,0.944],这与识别或结果相关的泄漏风险一致。在相同二元子集中的预注册分解消融发现,将相同输入转换为类型化阶段路线与全数据包分数无显著差异(Delta AUC = -0.144,95% CI [-0.471,0.176]),与确定性启发式方法也无显著差异(Delta AUC = -0.089,95% CI [-0.412,0.278])。该试点建立了可审计的可行性结果并揭示了失败模式;但未建立前瞻性校准、因果决策改进、路线分解优势或跨域有效性。
英文摘要
Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop. We study the complementary setting in which ground truth is delayed, censored, or private, so deterministic code cannot check correctness at scoring time and must instead issue a code-owned provisional forecast. RouteCast instantiates this regime for model-generated typed strategic routes: models propose candidate routes and structured factors; point-in-time evidence, reference classes, and deterministic transformations produce a provisional forecast-ranking; later outcomes evaluate the forecast. In a retrospective venture pilot on 21 binary-outcome cases (6 positive, 15 negative), the whole-packet RouteCast score showed preliminary retrospective discrimination (AUC 0.756, 95% CI [0.471,0.980]), while a blind LLM judge reached AUC 0.678 [0.419,0.897] and an identity-exposed LLM judge reached AUC 0.761 [0.515,0.944], consistent with recognition- or outcome-related leakage risk. A preregistered decomposition ablation on the same binary subset found that converting the identical inputs into typed staged routes was indistinguishable from the whole-packet score (Delta AUC = -0.144, 95% CI [-0.471,0.176]) and from a deterministic heuristic (Delta AUC = -0.089, 95% CI [-0.412,0.278]). The pilot establishes an auditable feasibility result and exposes failure modes; it does not establish prospective calibration, causal decision improvement, route-decomposition advantage, or cross-domain validity.
Comments11 pages, 2 figures