发表机构
MyndwareMed(迈恩德韦尔医疗)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文区分生成式上下文与受控状态,定义可问责临床AI的审计标准与信息要求,提出六级成熟度框架,将当前LLM临床实践定位于高能力低成熟度,为可问责临床AI提供概念与审计工具。
AI 中文摘要
大型语言模型(LLMs)已成为临床人工智能的主流交互界面,但其提供的“文本输入、文本输出、每次一个上下文窗口”的交互方式,并未对患者当前的真实状态维护明确、持久且受控的表征。本文提出,纵向临床推理是部分可观测下的状态估计问题,临床AI成败的核心并非模型读取病历的流畅性,而是其推理所依据的患者状态的受控性。我们区分生成式上下文与受控状态,将临床AI常混淆的五个对象(真实状态、观测值、证据、信念、模拟状态)加以分离,定义了一套分层的治理标准,可用于审计任何临床AI系统;并证明可问责性的操作定义可分解为四项信息要求:具备感知时间版本控制的不可变证据账本、与累积证据相区分的信念状态、观测过程模型,以及声明级因果类型。我们明确指出,这种分解是分析性的而非必然性定理,其价值在于概念层面的梳理:将“可问责临床AI”从口号转化为审计工具。一套六级成熟度框架区分了系统可管控的内容与可计算的内容,将当前以LLM为核心的实践定位为高能力但低成熟度。本文完全独立自足:框架提出的四个研究问题在引言中阐明,结论记录了本文针对每个问题所取得的进展;未来工作将开发该架构的可构建核心,以及通向完整临床世界模型(Clinical World Models)的研究计划。本文未声明任何实证结果。
英文摘要
Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about a patient. This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over. We distinguish generated context from governed state; separate five objects that clinical AI habitually conflates (true state, observations, evidence, belief, and simulated state); define a tiered governance standard against which any clinical AI system can be audited; and show that an operational definition of accountability decomposes into four information requirements: an immutable evidence ledger with awareness-time versioning, a belief state distinct from accumulated evidence, an observation-process model, and claim-level causal typing. We are explicit that this decomposition is analytic rather than a necessity theorem, and that its value is conceptual hygiene: it converts "accountable clinical AI" from a slogan into an audit instrument. A six-level maturity framework separates what a system makes governable from what it can compute, locating current LLM-centric practice at high capability but low maturity. The paper is fully self-contained: the four research questions the framework poses are stated in the introduction, and the conclusion records what the paper establishes toward each; future work develops the buildable core of the architecture and the research program toward full Clinical World Models. No empirical result is claimed here.