发表机构
Oracle Health and Life Sciences(甲骨文健康与生命科学部门)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对合成临床基准的结构不真实问题,提出在不破坏下游效用检查的前提下提升其真实性的方法,通过确定性修正改善基准指标,明确需将效用作为约束而非真实性的充分证据。
AI 中文摘要
面向企业AI智能体的合成临床基准虽可通过现有效用检查,仍可能存在结构不真实的问题,尤其在隐私敏感、难以获取运营数据的医疗场景中。本文研究如何在不破坏实际使用的下游效用检查的前提下改进这类基准,将基准修正定义为效用约束下的真实性提升:数据集的修改需在保持高于运营效用下限的同时提升真实性。本文在基于Synthea生成患者、经演示电子健康记录工作流程处理并与运营数据采用相同下游流程的护理差距基准上实例化该思路,真实性通过缺失结构、简洁性、结构合理性及人群对齐度衡量。基线基准极为单薄:采样对缺失率达79.44%,仅12.75%的行可操作,38.94%的患者无任何可操作指标,前三令牌集中度达100.0%。两种确定性修正可在保持高于当前效用下限的同时改善这些指标,而朴素的 densification 对照组则保留了不真实的模板化特征。本文进一步表明,基准内部真实性与对聚合运营参考的源保真度是相关但不同的目标。这些结果表明,应明确优化合成基准质量,将效用视为约束之一而非真实性的充分证据。
英文摘要
Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. We study how to improve such benchmarks without breaking the downstream utility checks already used in practice. We formulate benchmark revision as utility-constrained realism improvement: dataset changes should increase realism while staying above an operational utility floor. We instantiate this idea on a care-gap benchmark derived from Synthea-generated patients exercised through demonstration electronic health record workflows and then processed by the same downstream pipeline as operational data. Realism is measured through missingness structure, simplicity, structural plausibility, and population alignment. The baseline benchmark is extremely thin: sampled-pair missingness is 79.44%, only 12.75% of rows are actionable, 38.94% of patients have zero actionable measures, and top-three token concentration reaches 100.0%. Two deterministic revisions improve these panels while remaining above the current utility floor, whereas a naive densification control preserves unrealistic templating. We further show that internal benchmark realism and source fidelity to an aggregate operational reference are related but distinct objectives. These results suggest that synthetic benchmark quality should be optimized explicitly, with utility treated as one constraint rather than as sufficient evidence of realism.