发表机构
Ben-Gurion University of the Negev; University of Cambridge; Clalit Health Services; Tel Aviv University; School of Brain Sciences and Cognition, Ben-Gurion University of the Negev; Azrieli National Center for Autism and Neurodevelopment Research(内盖夫本-古里安大学; 剑桥大学; 克拉利特医疗服务机构; 特拉维夫大学; 内盖夫本-古里安大学脑科学与认知学院; 阿兹里利国家自闭症与神经发育研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出基于理论的 FORMA 框架,将认知模型编译为图来生成符合临床结构的创伤后应激障碍案例 vignettes,经评估其质量高且人口统计学差异小,可作为合成临床文本生成的可审计规范。
AI 中文摘要
大型语言模型可生成流畅的临床案例 vignettes,但仅流畅性无法确保其符合可明确指定的临床结构。我们推出 FORMA,这是一个基于理论的框架,它将某一障碍的认知模型编译为有向加权图,对该图的特定于个体的配置进行采样,并验证生成的 vignette 是否保留了指定的组件和因果关系。我们使用 FORMA 基于 Ehlers 和 Clark 的认知模型针对创伤后应激障碍进行实例化,在 500 个 persona、11 个生成模型和 3 个消融条件下生成了 16500 个 vignettes。评估结合了外部边恢复探针、两名临床专家、一个规模化 LLM 评判器以及一项有 100 名持照从业者参与的临床医生用户研究。从全条件 vignettes 中可恢复认知图(MCC = +0.41,AUC = 0.70),但无法从零样本生成中恢复(MCC = +0.01,AUC = 0.50)。专家对全条件 vignettes 的评分远高于零样本替代方案,且临床医生认为全条件 vignettes 有 85% 的概率是人工撰写的,而零样本的这一比例为 22%。FORMA 还将感知质量方面的人口统计学差异降低了 1.5 至 7 倍。这些结果表明,认知构想可作为可审计的规范,用于可扩展的合成临床文本生成。包含数据和代码的存储库可在线获取:this https URL。
英文摘要
Large language models can generate fluent clinical case vignettes, but fluency alone does not ensure fidelity to a specifiable clinical structure. We introduce FORMA, a theory-grounded framework that compiles a cognitive model of a disorder into a directed weighted graph, samples a person-specific configuration of that graph, and validates whether the generated vignette preserves the specified components and causal links. We instantiate FORMA on Posttraumatic Stress Disorder using the Ehlers and Clark cognitive model, generating 16,500 vignettes across 500 personas, 11 generation models, and three ablation conditions. Evaluation combines an external edge-recovery probe, two clinical experts, a scaled LLM judge, and a clinician user study with 100 licensed practitioners. The cognitive graph is recoverable from full-condition vignettes (MCC = +0.41, AUC = 0.70) but not from zero-shot generation (MCC = +0.01, AUC = 0.50). Experts rate full vignettes substantially higher than zero-shot alternatives, and clinicians perceive them to be human-written 85% of the time, compared with 22% for zero-shot. FORMA also reduces demographic disparity in perceived quality by 1.5-7x. These results show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation. A repository with the data and code is available online: https://github.com/Amit-Oren/FORMA.
CommentsAccepted to Findings of the Association for Computational Linguistics: EMNLP 2026. Code and data: https://github.com/Amit-Oren/FORMA