类型化决策模型:早期证据审计与评估清单
Typed Decision Models: An Early Evidence Audit and Evaluation Checklist
浏览论文内容
中文总结 AI 辅助
本文审计了类型化决策模型(TDM)的早期证据,基于28篇论文提出14项评估清单,指出其优势在延迟与成本,准确性仍存差距。
中文摘要 AI 辅助
类型化决策模型(TDM)返回调用方定义选项上的概率分布,而不生成文本。TypeSafe于2026年9月15日发布了商业类型化决策模型Jev,随后数日内出现了一小批评估和复现工作。我们回顾了9月19日至24日期间发布的28篇论文,并将其发现与早期关于标签概率分类、约束解码、重排序、校准和模型级联的研究联系起来。在这早期文献中,类型化读出本身并未显示出相对于可比标签概率读出的独立准确性优势。Jev最明显的优势在于延迟和成本,而在较难任务上仍存在准确性差距。在实际部署中,置信度常被用于决定何时将任务交给更强的模型或人类(弃权(不执行))。我们利用这些研究中反复出现的弱点,为未来TDM工作推导出一份14项评估清单。由于证据仅覆盖单一托管模型发布后前九天的数据,本综述应被视为早期证据图谱,而非对该模型类别的定论性评估。
英文摘要
Typed decision models (TDMs) return probability distributions over caller-defined options without generating text. TypeSafe released Jev, a commercial typed decision model, on 15 September 2026, and a small body of evaluation and replication work appeared within days. We review 28 papers posted between 19 and 24 September and relate their findings to earlier work on label-probability classification, constrained decoding, reranking, calibration, and model cascades. In this early literature, the typed readout itself has not shown an independent accuracy advantage over comparable label-probability readouts. Jev's clearest gains are in latency and cost, while accuracy gaps remain on harder tasks. In practical deployments, confidence is often used to decide when to defer to a stronger model or a human. We use recurring weaknesses in these studies to derive a 14-item evaluation checklist for future TDM work. Because the evidence covers only the first nine days after the release of one hosted model, the review should be read as an early evidence map rather than a settled assessment of the model class.