发表机构
University of Memphis(孟菲斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对临床诊断智能体的自主停止可靠性问题,提出风险约束停止层Cros,结合错误排序与精确检验,在腹痛基准上显著降低选择性错误与成本,但结果仅为探索性证据。
AI 中文摘要
临床诊断智能体不仅必须决定下一步请求哪项检查,还必须决定何时做出诊断或推迟诊断。现有的智能体基准主要评估固定或不受约束交互后的准确性,使得自主停止的可靠性隐式化。我们提出了Cros,一个风险约束的停止层,结合了逐状态错误排序、在不相交的开发数据划分上的策略设计,以及针对完整序贯策略的选择性诊断错误和最小自主覆盖率的LTT式精确检验。其有限样本保证要求在访问校准标签之前冻结候选族、检验规则以及任何随机化。在一个包含1,834个回合的基于MIMIC的腹痛基准上,完整的排序器实现了探索性状态错误AUROC 0.853,而最大类别概率为0.715,骨干网络的原生停止得分为0.552。在先前见过的367个回合的评估划分上,对冻结的Cros权重进行解析平均,在78.8%的覆盖率下实现了16.9%的选择性错误,成本为5.57,测试次数为0.68,而在原生停止下,在100%覆盖率下错误率为30.8%,成本为8.14,测试次数为1.53。强制继续是非单调的:仅使用HPI时错误率为28.3%,完整检查后为34.3%。然而,在这个见过的划分上,均匀权重混合消融更便宜,尽管错过了锁定的开发边界,并且Cros在20次开发重划分中仅在6次中名义上满足联合标准。由于评估标签在早期开发期间已被检查,这些发现提供了探索性可行性和审计证据,而非确证性安全证书。
英文摘要
Clinical diagnosis agents must decide not only what test to request next, but also when to diagnose or defer. Existing agent benchmarks largely evaluate accuracy after fixed or unconstrained interaction, leaving autonomous stopping reliability implicit. We present Cros, a risk-constrained stopping layer combining state-wise error ranking, policy design on disjoint development splits, and LTT-style exact tests of selective diagnostic error and minimum autonomous coverage for complete sequential policies. Its finite-sample guarantee requires the candidate family, testing rule, and any randomization to be frozen before calibration labels are accessed. On a 1,834-episode MIMIC-derived abdominal-pain benchmark, the full ranker achieves exploratory state-error AUROC 0.853, compared with 0.715 for maximum class probability and 0.552 for the backbone's native stop score. On the previously viewed 367-episode evaluation split, analytically averaging over the frozen Cros weights yields 16.9% selective error at 78.8% coverage, cost 5.57, and 0.68 tests, versus 30.8% error at 100% coverage, cost 8.14, and 1.53 tests under native stopping. Forced continuation is non-monotone: error is 28.3% with HPI alone and 34.3% after full workup. However, the uniform-weight mixture ablation is cheaper on this viewed split despite missing the locked development margins, and Cros nominally satisfies the joint criterion in only 6 of 20 development resplits. Because evaluation labels were inspected during earlier development, these findings provide exploratory feasibility and audit evidence, not a confirmatory safety certificate.